Chinese company PsiBot has unveiled Psi-R2.5 — a foundation model for robots that is shown a new manipulation task once, instead of being fed dozens of teleoperated repetitions. The key is not a bigger model but how the data is produced: the company reversed its own pipeline to mass-produce human–robot pairs that correspond frame by frame.
Key takeaways
- Two-layer architecture: a planner built on QwenVL3.5-4B, a controller built on Wan2.2-TI2V-5B
- “Strong pair data”: frame-by-frame correspondence and an action replayable directly on the robot
- Phone-box assembly: 99% success after a few iterations, completed in one to two working days
- Evaluation set: 50 complex real-robot tasks with randomised initial states
- The pipeline also works on footage from an ordinary phone
What makes pair data “strong”
Until now, human and robot data pairs were matched only by task semantics — a person and a robot picked up the same bottle, but everything else differed. In the strong variant, the image outside the acting body is essentially identical, and the recorded action can be replayed directly on the robot. The company describes this as pulling the dynamics of a human hand into the same domain?Data domain: The shared properties of a recording — framing, lighting, motion dynamics. Data from different domains is treated by the model as two separate worlds. as the robot’s.
| Criterion | Weak pair data | Strong pair data |
|---|---|---|
| What matches | task semantics only | the image outside the acting body |
| Temporal alignment | none | frame by frame |
| Action on the robot | not directly replayable | replayable directly |
The pipeline runs backwards
PsiBot previously went from human data to robot data using its Psi-W0 world model and reinforcement learning — expensive, with rollouts for every sample. The direction is now reversed.
The diagram contrasts two directions of the same pipeline. In the old one, human footage was the starting point and every sample needed a costly rollout. In the new one, real robot data comes first and the world model generates the matching human-hand footage for it.
Three models in the puzzle:
Two tests check quality: the data must replay on a physical robot, and it must support post-training that generalises to related tasks.
What was not disclosed
Psi-R2.5 has no paper and no repository — the company blog is the only technical source. Numerical results from the 50-task set were not published, and “one demonstration instead of fifty” is the article’s framing, not a controlled comparison.
Why it matters
For two years embodied AI has scaled on hours of recorded footage. PsiBot argues that beyond 100,000 hours of human footage, adding more stops helping, because the signal-to-noise ratio in camera footage is far worse than in data from a real robot. If that holds, the advantage will go to teams with a better conversion pipeline, not a bigger disk.
What’s next
- The company has promised technical details of its in-context learning in a later publication
- Numerical results from the 50-task evaluation set remain unpublished
- The absence of a peer-reviewed paper and code makes independent verification hard
Sources
- 机器之心 — 具身智能开始拼高质量数据了!灵初Psi-R2.5用强 Pair Data重做训练链路
- PsiBot — Scaling Pair Data for Embodied Intelligence
- PsiBot — About us





