RFI
How it works
RFI is implemented through three architectural choices. (1) Dual action experts: a reactive action expert is grafted beside the original action expert; it decodes actions causally within an action chunk and consumes the most recent 6D end-effector force. (2) Causal force memory: force enters via a causal force memory and zero-initialized cross attention, so force history is incorporated without modifying the initial action output; the memory is latency-aligned with the robot state. (3) Within-chunk refresh: the slow vision-language prefix is computed once and cached, and each within-chunk refresh re-encodes the latest latency-aligned force history without another full forward pass through the vision-language module. Prior preservation comes from initialization: the reactive expert copies the original action weights and uses shifted causal attention so force is added as a zero-initialized residual, meaning that before post-training the force residual is exactly zero and the augmented model starts from the same action output as the base VLA. Training uses an additive flow-matching objective that jointly supervises both streams: vision-only batches mask the force memory while online correction batches activate it (a fixed 1:1 offline:online sampling ratio). A repeated online DAgger loop collects force-enabled human corrections (through Flexiv TDK) and mixes them with offline data to address policy-dependent distribution shift.
Problem solved
Pretrained VLA policies are largely vision-driven and fail in contact-rich manipulation when the scene is occluded, depth is ambiguous, and small force errors push execution off the offline demonstration distribution. Force reveals contact states vision alone misses (whether cloth was grasped, a book bottomed out, or a ring jammed). RFI adds this reactivity without discarding the base model's knowledge.
Components
A second action expert grafted beside the original action expert of the base VLA. It decodes actions causally within an action chunk and consumes the most recent 6D end-effector force, forming a parallel reactive stream without replacing the base policy.
The channel through which force signals enter the policy via zero-initialized cross attention. It incorporates force history without modifying the initial action output and provides a latency-aligned force memory synchronized with the actual robot state.
An initialization mechanism: the reactive expert copies the original action weights and uses shifted causal attention so force is added as a zero-initialized residual. Before post-training the force residual is exactly zero, so the augmented model starts from the same action output as the base VLA.
During execution the slow vision-language prefix is computed once and cached. Each within-chunk refresh re-encodes the latest latency-aligned force history without another full forward pass through the vision-language module.
An additive flow-matching objective jointly supervises the base and reactive streams. Vision-only batches mask the force memory while online correction batches activate it, allowing one model to train on both dataset types without synthetic labels; a fixed 1:1 offline:online sampling ratio is used.
A repeated online DAgger loop collects force-enabled human corrections (through Flexiv TDK) and mixes them with offline data to address policy-dependent distribution shift.
Official
Implementation
Naively adding a force pathway can disturb the pretrained VLA output and erase its general manipulation knowledge.
If force history is not synchronized with the actual robot state, reactive actions respond to stale contact.
Small force errors push execution off the offline demonstration distribution, which offline data alone does not cover.
Evolution
The paper "Never Too Late for Force" (arXiv:2607.14236, CoRL 2026) introduces Reactive Force Injection as the core of the LIFT framework for force-aware VLA post-training.
Hyperparameters (configurable axes)
Mixing proportion of offline data and online corrections during post-training. The paper uses a fixed 1:1 value.
Span of 6D force history fed to the reactive expert through the causal force memory; must be latency-aligned with the robot state.
Number of actions decoded per chunk; determines how often the within-chunk refresh with the latest force history occurs.
Execution paradigm
Parallelism
The slow vision-language prefix is computed once and cached, while fast within-chunk force refreshes run without a full vision-language forward pass.