DR
A 23B World Action Model built on Wan2.1 video diffusion that jointly predicts future world frames and robot actions, giving 2x better zero-shot generalization than leading VLAs.
🔬 Research🔬 Research only⚖ Open weightsWorld ModelVision-Language-Action modelRobotics foundation model
Parameters
23B
parameters
Release date
17 February 2026
Access:DownloadDeployment:💻 Local
Overview
Classification
World ModelVision-Language-Action modelRobotics foundation model
Access & deployment
Download
Local
Weights: Open weights
Key parameters
🧩 Parameters: 23B
✓ Fine-tuning
📥 Input: video, image, text, robot state data
Technical specification
Parameters
23B
parameters
License
CC-BY-NC-4.0
Hardware requirements
23B BF16 weights (safetensors) on the Wan2.1-I2V video diffusion backbone; ~7Hz closed-loop control after optimizations.
Features:✓ Fine-tuning
Modalities
⬇ Input
videoimagetextrobot_state_data
⬆ Output
videorobot_actionsmotion_trajectories
Capabilities and applications
Native model capabilities
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
World simulation
Model's ability to generate coherent, interactive simulations of physical environments — maintaining geometry, lighting, and physics during exploration.
Category: multimodal
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Zero-shot learning
The model's ability to perform a new task without dataset-specific training or hyperparameter tuning — prediction is produced in a single pass from context.
Category: other
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Benchmark results
2 benchmarks
Generalizacja zero-shot vs SOTA VLA
zero-shot
2×
📄 paper
2x improvement in zero-shot generalization over leading VLAs per the paper (2026).
Sterowanie w pętli zamkniętej
~7 Hz
📄 paper
Closed-loop control frequency after optimizations.
Technical architecture
Core Architecture
Training Techniques