Sim-to-Real
How it works
A typical sim-to-real pipeline: (1) build a physics and rendering simulator mirroring the robot and task; (2) train the policy in simulation, usually with reinforcement learning across many parallel instances; (3) close the reality gap with one or more methods: domain randomization (randomizing textures, masses, friction, latencies), domain adaptation (tuning representations toward real data), system identification (calibrating simulator parameters to measurements from the real robot); (4) deploy and evaluate on the physical robot; (5) optionally fine-tune on a small amount of real data. The automatic domain randomization (ADR) variant gradually increases difficulty as the policy improves.
Problem solved
Training robots directly on physical hardware is slow, costly and risky (wear, failures, safety). Simulation solves this by providing unlimited data, but policies learned in simulation usually fail on a real robot because of the reality gap. Sim-to-Real closes that gap, letting skills transfer from simulation to hardware.
Components
Randomly varying simulation parameters (textures, lighting, masses, friction, latencies) so that reality appears as just another training variation.
Tuning the representation or policy so that the feature distribution from simulation approaches that of real-world data.
Calibrating simulator parameters (dynamics, friction, latencies) to measurements from the real robot so the simulation better matches hardware.
A physics and rendering engine (often massively parallel on GPU) generating training data that mirrors the robot and task.
Implementation
Unmodeled physics, latencies and sensor noise cause a policy that excels in simulation to fail on hardware.
Too-wide randomization yields overly conservative, suboptimal policies in the real world.
Evolution
Sadeghi and Levine demonstrate collision-free flight learned purely from randomly rendered simulation images, without a single real training image.
Tobin et al. formalise domain randomization, enabling networks trained purely on synthetic data to transfer to real hardware.
OpenAI trains an in-hand manipulation policy in simulation with domain randomization and transfers it to a physical Shadow robotic hand.
OpenAI solves the Rubik's cube one-handed using automatic domain randomization (ADR) of increasing difficulty.
The survey by Zhao, Peña Queralta and Westerlund organises the main approaches: domain randomization and adaptation, imitation learning, meta-learning and knowledge distillation.
NVIDIA Isaac Gym and locomotion work (e.g. the ANYmal quadruped) demonstrate training thousands of parallel environments on GPU and transferring walking policies to real robots in minutes.
Hyperparameters (configurable axes)
The span and distribution of randomized parameters (textures, lighting, masses, friction, latencies). Too-narrow ranges fail to close the reality gap; too-wide ranges yield overly conservative policies.
How many simulation instances run at once on the GPU. Directly affects data throughput and training speed.
The trade-off between physics/rendering accuracy and steps per second. Higher fidelity narrows the reality gap at the cost of speed.
The rate at which randomization ranges are widened as the policy improves. Central to OpenAI's approach (Rubik's cube).
The amount of physical-robot data used to optionally fine-tune the policy after transfer. Often minimal or zero (zero-shot).
Compute bottleneck
The bottleneck is the rate at which the simulator generates data — physics stepping and (optionally) rendering. Physics fidelity, the number of parallel environments and rendering resolution directly limit how fast reinforcement-learning training data is produced.
Parallelism
The biggest advantage of sim-to-real: GPU simulators (e.g. NVIDIA Isaac Gym / Isaac Lab) run thousands of parallel environments at once, giving enormous data throughput and cutting policy training from days to minutes. Deployment (inference) on the physical robot, by contrast, is sequential — the policy runs in real time inside the control loop.
Hardware requirements
Massively parallel physics simulation and reinforcement-learning policy training (e.g. Isaac Gym / Isaac Lab) run on the GPU; thousands of simultaneous environments are the domain of graphics cards.