Robots Atlas>ROBOTS ATLAS
WALL-WM

WALL-WM

WALL-WM · Family: WALL
X Square Robot’s open-source (Apache-2.0) world action model (WALL series): event-grounded VLA pretraining, video and action prediction, Muon optimizer.
✓ Active✓ Public access⚖ Open sourceWorld ModelRobotics foundation modelVision-Language-Action model📁 WALL
Release date
29 May 2026
Access:DownloadDeployment:💻 Local📱 On-device

Overview

WALL-WM is an open-source world action model for general-purpose embodied AI, developed by X Square Robot (自变量机器人, Shenzhen) and released on 29 May 2026. It belongs to the WALL series and is described in the paper “WALL-WM: Carving World Action Modeling at the Event Joints” (arXiv:2606.01955).

Its key idea is a shift from chunk-centric learning over fixed time windows to event-grounded vision-language-action pretraining, where the atomic unit of learning is a semantically coherent action event (e.g. grasping, lifting). This lets the model learn task objectives rather than memorizing pixel sequences, and better reconciles language semantics, continuous visual dynamics and control-level actions.

The model couples video prediction with action prediction at event boundaries. It offers two inference modes: an event mode (variable-length) and a unified mode with Staircase Decoding. It uses a Muon-optimizer-based pretraining infrastructure and a three-layer architecture.

WALL-WM is released under the Apache-2.0 license together with code and tooling (GitHub wall-x, Hugging Face), permitting commercial and non-commercial use, self-hosting and fine-tuning. It complements the WALL series with world-modeling and action-consequence prediction for planning and robot policy learning.

Classification
World ModelRobotics foundation modelVision-Language-Action model
Family: WALL
Access & deployment
Download
LocalOn-device
Weights: Open source
Key parameters
✓ Fine-tuning
📥 Input: text, image, video, robot state data
Robotics
Environment modelingSpatial predictionMotion planningEmbodied task planningRobot manipulation

Technical specification

License
Apache-2.0
Hardware requirements
Open-source model deployed locally/on-device; training and inference require GPU acceleration.
Features:Fine-tuning
Modalities
⬇ Input
textimagevideorobot_state_data
⬆ Output
videorobot_actionsmotion_trajectories

Capabilities and applications

Native model capabilities
World simulation
Model's ability to generate coherent, interactive simulations of physical environments — maintaining geometry, lighting, and physics during exploration.
Category: multimodal
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Planning
Forming and executing action plans for complex tasks.
Category: planning
Multimodal understanding
Category: multimodal
Robotics
Environment modelingSpatial predictionMotion planningEmbodied task planningRobot manipulation

Deployment and security

☁ Available on platforms