Robots Atlas>ROBOTS ATLAS
GAIA-2

GAIA-2

GAIA-2 · Family: GAIA
Wayve’s controllable multi-view world model for autonomous driving (latent diffusion). Generates realistic video for simulation and training (UK/US/DE).
✓ Active🔬 Research onlyWorld ModelVideo generationMultimodal📁 GAIA
Release date
26 March 2025
Deployment:☁ Cloud

Overview

GAIA-2 is a controllable, multi-view generative world model developed by Wayve for autonomous driving, introduced on 26 March 2025. It is described in the paper “GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving” (arXiv:2503.20523). The model generates realistic, temporally and spatially consistent multi-camera video for simulation, scenario testing and augmenting training data.

Its architecture consists of a video tokenizer that compresses sequences into a continuous latent space and a latent-diffusion world model that predicts future scene states. Unlike GAIA-1 (autoregressive), GAIA-2 uses latent diffusion, natively supports multiple cameras and eliminates temporal discontinuities, improving motion smoothness and cross-view coherence.

The model accepts rich structured conditioning: ego-vehicle actions (speed, steering curvature), dynamic agent behaviors (from 3D bounding boxes), road attributes (number of lanes, speed limits, crossings, traffic lights, intersections), environmental factors (weather, time of day) and CLIP plus proprietary model embeddings. It covers data from the UK, US and Germany.

GAIA-2 is used to generate synthetic scenarios (including rare and safety-critical ones), test out-of-distribution generalization and augment real-world data. It is a research, proprietary model (weights not publicly available).

Classification
World ModelVideo generationMultimodal
Family: GAIA
Access & deployment
Cloud
Weights: Closed
Key parameters
📥 Input: video, text, structured data

Technical specification

License
Proprietary (Wayve)
Modalities
⬇ Input
videotextstructured_data
⬆ Output
video

Capabilities and applications

Native model capabilities
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Synthetic data generation
Generating synthetic datasets that preserve the statistical properties of the original — used for model training, testing, and privacy protection.
Category: structured_generation
World simulation
Model's ability to generate coherent, interactive simulations of physical environments — maintaining geometry, lighting, and physics during exploration.
Category: multimodal