Shanghai Jiao Tong University researchers have presented LIFT, a method that teaches a pretrained VLA policy to react to contact force without a single force label in pretraining. It was accepted at CoRL 2026, code publicly available.
Key takeaways
- LIFT grafts a reactive action expert beside the original inside a pretrained pi-0.5 policy
- Recent 6D force history enters via zero-initialised cross attention, so the model starts out identical to the base
- Tested on a Flexiv Rizon 4S: towel folding, book insertion, Hanoi ring placement
- Evaluation: 30 autonomous rollouts per checkpoint, 95 percent confidence intervals
- Corresponding authors: Chuan Wen and Cewu Lu, arXiv 2607.14236
Force arrives last, the prior stays intact
Pretrained VLA policies understand language but steer almost entirely on images. In contact that fails: the scene can be occluded, depth ambiguous, and a small force error pushes execution outside the demonstration distribution.
LIFT copies the original action expert weights into a new reactive stream and wires the force pathway through zero-initialised cross attention, so the force residual starts at exactly zero and the model matches the base.
The reactive stream then decodes actions causally inside the chunk?action chunking: Predicting a whole short sequence of actions at once instead of a single control step. A chunk is one such batch of actions., and the slow vision-language prefix is cached after one pass.
The diagram shows why LIFT does not break the base model. The zero-initialised output projection keeps the force residual at exactly zero before post-training, so the policy behaves exactly like the pretrained π0.5. The force signal only starts shaping actions once human corrections arrive on the robot.
Zero force labels in vision data
One model trains on two inconsistent datasets. Handheld demonstrations have their force memory masked before cross attention, blocking gradients in the force pathway. Force readings appear only in human corrections gathered on the robot through Flexiv TDK, and offline-to-online sampling holds a 1:1 ratio.
What the ablations show
Force-enabled post-training learns faster and reaches higher peak and final performance than vision-only online DAgger on all three tasks. The authors report no percentages — comparisons rest on learning curves.
The ablations also show what fails: a frozen correction buffer instead of repeated online aggregation scores zero on book insertion, and a lightweight residual policy over a frozen π0.5 falls far behind.
Why it matters
The dominant approach to force in robotics assumes collecting force data from the start, which rules out ready-made vision-only checkpoints. LIFT reverses the order: it takes a pretrained VLA and adds the modality at the end.
That changes data economics — what counts is a few dozen quality corrective episodes, not large force corpora. The entry barrier for contact-rich manipulation drops to hours of work at the bench.
What next
- Human correction throughput is the stated bottleneck — one online experiment takes 2–3 hours and about 20–30 episodes
- Evaluation covered one robot arm, so next come tests on other manipulators, force sensors, and end effectors
Sources
- Jiqizhixin — CoRL 2026 | 0条带力数据预训练!上海交大卢策吾、汶川团队让VLA学会力
- arXiv — Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection
- LIFT project page — LIFT: Never Too Late for Force
- GitHub — y-wng/lift





