Human recordings and robot executions are collected for the same task description, independently of one another — at different times, often in different scenes and at different speeds. The result is a dataset where correspondence exists at the task-label level but not at the frame or state level. A model trained on such data must infer the visual mapping (human hand vs gripper) and the kinematic one (a different motion profile) by itself, which is the source of the embodiment gap. PsiBot points out that moving from this regime to strongly paired data — rather than further increasing the number of recorded hours — was the key change in the Psi-R2.5 training pipeline.
Weak pair data does not itself solve a problem — it is the cheap default way of collecting human–robot data. The concept solves a terminological problem: it lets you state precisely what such data lacks (temporal alignment and scene consistency) and why the sheer scale of human footage does not translate into robot policy quality.
The human recording and the robot execution concern the same task in the sense of its description, with no further requirements.
There is no frame-to-frame or state-to-state correspondence between the demonstration and the execution.
Increasing the number of hours of weakly paired footage does not remove the embodiment gap and yields diminishing returns in policy quality.
A dataset labelled "paired" is often in fact weakly paired, if temporal alignment, scene consistency and action replayability were never checked.
PsiBot defined weak pair data in the Psi-R2.5 technical blog as the contrast to strong pair data.