Robots Atlas>ROBOTS ATLAS
Robotics

Sim-to-Real

2017ActivePublished: 24 August 2026Updated: 20 September 2026Published
Key innovation
Enabling cheap, safe and massively parallel training of control policies in simulation and then transferring them to a physical robot — by explicitly closing the 'reality gap', notably via domain randomization that treats reality as just another simulation variation.
Category
Robotics
Abstraction level
Pattern
Operation level
TrainingRobot controlModel
Use cases
Learning locomotion for legged robots (e.g. quadrupeds, humanoids)Dexterous manipulation and grasping (e.g. robotic hands)Navigation and collision avoidance for drones and mobile robotsVisual perception trained on synthetic dataMassively parallel GPU-based policy training

How it works

A typical sim-to-real pipeline: (1) build a physics and rendering simulator mirroring the robot and task; (2) train the policy in simulation, usually with reinforcement learning across many parallel instances; (3) close the reality gap with one or more methods: domain randomization (randomizing textures, masses, friction, latencies), domain adaptation (tuning representations toward real data), system identification (calibrating simulator parameters to measurements from the real robot); (4) deploy and evaluate on the physical robot; (5) optionally fine-tune on a small amount of real data. The automatic domain randomization (ADR) variant gradually increases difficulty as the policy improves.

Problem solved

Training robots directly on physical hardware is slow, costly and risky (wear, failures, safety). Simulation solves this by providing unlimited data, but policies learned in simulation usually fail on a real robot because of the reality gap. Sim-to-Real closes that gap, letting skills transfer from simulation to hardware.

Components

Domain randomizationClosing the reality gap

Randomly varying simulation parameters (textures, lighting, masses, friction, latencies) so that reality appears as just another training variation.

Domain adaptationDistribution alignment

Tuning the representation or policy so that the feature distribution from simulation approaches that of real-world data.

System identificationSimulation fidelity

Calibrating simulator parameters (dynamics, friction, latencies) to measurements from the real robot so the simulation better matches hardware.

High-fidelity simulatorTraining-data source

A physics and rendering engine (often massively parallel on GPU) generating training data that mirrors the robot and task.

Implementation

Implementation pitfalls
The reality gapHigh

Unmodeled physics, latencies and sensor noise cause a policy that excels in simulation to fail on hardware.

Fix:System identification, randomizing dynamics and latencies, and validating on the real robot.
Over-randomizationMedium

Too-wide randomization yields overly conservative, suboptimal policies in the real world.

Fix:Automatic domain randomization (ADR) and choosing ranges grounded in real data.

Evolution

Original paper · 2017 · IROS 2017 · Josh Tobin
Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, Pieter Abbeel
2016
CAD2RL — early sim-to-real

Sadeghi and Levine demonstrate collision-free flight learned purely from randomly rendered simulation images, without a single real training image.

2017
Domain randomization
Inflection point

Tobin et al. formalise domain randomization, enabling networks trained purely on synthetic data to transfer to real hardware.

2018
OpenAI Dactyl — dexterous manipulation

OpenAI trains an in-hand manipulation policy in simulation with domain randomization and transfers it to a physical Shadow robotic hand.

2019
Rubik's cube and automatic domain randomization

OpenAI solves the Rubik's cube one-handed using automatic domain randomization (ADR) of increasing difficulty.

2020
Systematising sim-to-real methods

The survey by Zhao, Peña Queralta and Westerlund organises the main approaches: domain randomization and adaptation, imitation learning, meta-learning and knowledge distillation.

2021
Massively parallel GPU simulation
Inflection point

NVIDIA Isaac Gym and locomotion work (e.g. the ANYmal quadruped) demonstrate training thousands of parallel environments on GPU and transferring walking policies to real robots in minutes.

Hyperparameters (configurable axes)

Domain randomization rangesCritical

The span and distribution of randomized parameters (textures, lighting, masses, friction, latencies). Too-narrow ranges fail to close the reality gap; too-wide ranges yield overly conservative policies.

±20% masy i tarciaTypical dynamics randomization in locomotion.
Number of parallel environmentsHigh

How many simulation instances run at once on the GPU. Directly affects data throughput and training speed.

4096Order of magnitude typical for walking training on Isaac Gym.
Simulation fidelity vs speedHigh

The trade-off between physics/rendering accuracy and steps per second. Higher fidelity narrows the reality gap at the cost of speed.

Automatic domain randomization (ADR) scheduleMedium

The rate at which randomization ranges are widened as the policy improves. Central to OpenAI's approach (Rubik's cube).

Real-data fine-tuning budgetMedium

The amount of physical-robot data used to optionally fine-tune the policy after transfer. Often minimal or zero (zero-shot).

Compute bottleneck

Simulation throughput

The bottleneck is the rate at which the simulator generates data — physics stepping and (optionally) rendering. Physics fidelity, the number of parallel environments and rendering resolution directly limit how fast reinforcement-learning training data is produced.

Depends on
Wierność fizyki vs prędkośćLiczba równoległych środowisk

Parallelism

Parallelism level
Fully parallel

The biggest advantage of sim-to-real: GPU simulators (e.g. NVIDIA Isaac Gym / Isaac Lab) run thousands of parallel environments at once, giving enormous data throughput and cutting policy training from days to minutes. Deployment (inference) on the physical robot, by contrast, is sequential — the policy runs in real time inside the control loop.

Scope
Training

Hardware requirements

Primary

Massively parallel physics simulation and reinforcement-learning policy training (e.g. Isaac Gym / Isaac Lab) run on the GPU; thousands of simultaneous environments are the domain of graphics cards.