Robots Atlas>ROBOTS ATLAS
Psi-W0

Psi-W0

Psi-W0
PsiBot's action-conditioned world model: predicts future video from images, language and robot action trajectories; the second half of the dual-system architecture alongside Psi-R2.
✓ Active⏳ Limited accessWorld ModelRobotics foundation modelVideo generation
Release date
10 April 2026
Access:HostedDeployment:💻 Local☁ Cloud

Overview

Psi-W0 is an action-conditioned world model developed by the Chinese company PsiBot (灵初智能). It was formally released on 10 April 2026 together with the Psi-R2 model — on the same occasion the company open-sourced its first 1,000 hours of human hand operation data.

The model takes images, language and a robot action trajectory as input, and outputs predicted future video. Within PsiBot's dual-system architecture, Psi-W0 is complementary to the Psi-R2 operation-policy model: it evaluates and improves policy performance and closes the data flywheel together with Psi-R2.

In the next-generation Psi-R2.5 model, Psi-W0 gained an additional role: it is used to inversely generate human-hand data from robot trajectories, which makes it possible to build strong pair data without manually collecting synchronized human footage.

Classification
World ModelRobotics foundation modelVideo generation
Access & deployment
Hosted
LocalCloud
Weights: Closed
Key parameters
📥 Input: image, text, robot state data

Technical specification

Modalities
⬇ Input
imagetextrobot_state_data
⬆ Output
video

Capabilities and applications

Native model capabilities
World simulation
Model's ability to generate coherent, interactive simulations of physical environments — maintaining geometry, lighting, and physics during exploration.
Category: multimodal
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Multimodal understanding
Category: multimodal
Synthetic data generation
Generating synthetic datasets that preserve the statistical properties of the original — used for model training, testing, and privacy protection.
Category: structured_generation
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video