Robots Atlas>ROBOTS ATLAS
GEN-1.5
AI Modelsโ€บGEN

GEN-1.5

1.5ย ยทย Family: GEN
Generalist AI's embodied foundation model (VLA) that learns new physical tasks one-shot from a single demonstration and outputs 100 Hz action trajectories.
โœ“ ActiveRobotics foundation modelVision-Language-Action modelMultimodal๐Ÿ“ GEN
Context window
~30 s wideo
tokens
Release date
19 August 2026

Overview

GEN-1.5 is an embodied foundation model (Vision-Language-Action) developed by Generalist AI, introduced on August 19, 2026. It processes video input (a context window of roughly 30 seconds) together with sensor data, proprioceptive (robot-state) data and language, and outputs action trajectories at 100 Hz.

According to the company, the model shows the beginnings of one-shot physical-skill learning: it can acquire a new task from a single 3โ€“12 second demonstration, and adapt few-shot via 1โ€“10 gradient steps on 1โ€“5 minutes of data. Reported results are an average 59% (ยฑ10%) success with one-shot in-context prompting and 83% (ยฑ9%) after few-shot fine-tuning (10 gradient steps), with weight changes below 0.15%.

The model was trained via scaled pretraining on large physical-interaction datasets collected in homes, warehouses and factories, with a stated continuous pretraining run of 8+ months. GEN-1.5 belongs to the GEN model family (successor to GEN-1).

Classification
Robotics foundation modelVision-Language-Action modelMultimodal
Family: GEN
Access & deployment
Weights: Closed
Key parameters
๐Ÿ“ Context: ~30 s wideo
๐Ÿ“ฅ Input: video, robot sensors, robot state data, text
Robotics
Dexterous manipulationRobot manipulationRobot controlEmbodied task planningMotion planningObject affordance understanding

Technical specification

Context window
~30 s wideo
tokens
Modalities
โฌ‡ Input
videorobot_sensorsrobot_state_datatext
โฌ† Output
robot_actionsmotion_trajectories

Capabilities and applications

Native model capabilities
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Video understanding
The model's ability to analyse and interpret video content โ€” recognising actions, motion, events and relationships between objects over time.
Category: video
Zero-shot learning
The model's ability to perform a new task without dataset-specific training or hyperparameter tuning โ€” prediction is produced in a single pass from context.
Category: other
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Cross-embodiment transfer
The ability of a single model to control robots with different morphologies (humanoids, dual-arm rigs, mobile platforms) without training a separate model per platform. Intelligence is decoupled from embodiment, so the same policy runs on hardware with different kinematics and dynamics.
Category: robotics
Planning
Forming and executing action plans for complex tasks.
Category: planning
Robotics
Dexterous manipulationRobot manipulationRobot controlEmbodied task planningMotion planningObject affordance understanding
Application domains

Benchmark results

2 benchmarks
One-shot in-context task success
success rate ยท One-shot, in-context prompting
59% (ยฑ10%)%
๐Ÿ“„ Oficjalny blog Generalist AI (GEN-1.5)
Few-shot task success (10 gradient steps)
success rate ยท Few-shot, 10 gradient steps on 1โ€“5 min of data
83% (ยฑ9%)%
๐Ÿ“„ Oficjalny blog Generalist AI (GEN-1.5)

Technical architecture