
PsiBot's embodied AI foundation model: a two-layer architecture (QwenVL3.5-4B + Wan2.2-IT2V-5B) rebuilt around strong human–robot pair data.
✓ Active⏳ Limited accessVision-Language-Action modelRobotics foundation modelMultimodal
Access:HostedDeployment:💻 Local☁ Cloud
Overview
Classification
Vision-Language-Action modelRobotics foundation modelMultimodal
Access & deployment
Hosted
LocalCloud
Weights: Closed
Key parameters
📥 Input: text, image, video, robot sensors…
Technical specification
Modalities
⬇ Input
textimagevideorobot_sensorsrobot_state_data
⬆ Output
robot_actionsmotion_trajectoriesmanipulator_controltext
Capabilities and applications
Native model capabilities
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
Multimodal understanding
Category: multimodal
Planning
Forming and executing action plans for complex tasks.
Category: planning
Zero-shot learning
The model's ability to perform a new task without dataset-specific training or hyperparameter tuning — prediction is produced in a single pass from context.
Category: other
Few-shot learning
The model's ability to perform a new task from a handful of examples provided directly in the prompt, without any weight updates or fine-tuning.
Category: other
Sample efficiency
The ability to reach strong performance using far fewer environment interactions or training examples.
Category: other
Cross-embodiment transfer
The ability of a single model to control robots with different morphologies (humanoids, dual-arm rigs, mobile platforms) without training a separate model per platform. Intelligence is decoupled from embodiment, so the same policy runs on hardware with different kinematics and dynamics.
Category: robotics
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Image understanding
Analysing and interpreting the content of images.
Category: vision