Explainer
Sim-to-real: training robots in simulation for the real world
Sim-to-real transfer means training a robot controller in simulation and deploying it on real hardware. The difference between the two is called the "reality gap".[1] Randomizing the simulator helps bridge it, because a model that has seen enough variation treats the real world as one more variation.[2] The method trained walking robots that worked outdoors with no real-world training,[3] and NVIDIA and others now sell simulation and "world model" tools for robot learning.[4]
Training a robot in the physical world is slow and expensive. Simulation is fast and safe, but never quite matches reality. Sim-to-real research is about closing that “reality gap”.[1] This page explains the main technique and where it works best.
The reality gap
A simulator is a video-game-like world with physics. A robot can practise millions of times there without breaking anything. But simulated friction, lighting and motors are never exactly like real ones. A controller that works perfectly in simulation can fail on a real robot. Researchers call this mismatch the “reality gap”.[1]
Domain randomization, introduced for vision in 2017, randomizes rendering so that “the real world may appear to the model as just another variation”.[2] OpenAI’s automatic domain randomization grew the randomization range as the policy improved. It trained a robot hand in simulation that then solved a Rubik’s cube on real hardware.[5]
Where it works: walking
Sim-to-real works especially well for walking. A four-legged robot trained only in simulation walked over mud, snow and rubble outdoors without any extra training.[3] A human-sized humanoid was later taught to walk the same way.[6]
The 2020 ANYmal controller was trained with RL in simulation and showed zero-shot generalization to natural terrain.[3] A 2023 humanoid controller used a causal transformer over the history of proprioceptive observations and actions, trained with RL across randomized simulated environments. It adapted its behaviour in context on real terrain without weight updates.[6]
Where it works: hands and whole bodies
Simulation also works for hands. Researchers trained a robot hand in simulation to turn objects in its fingers and then used it on a real hand.[7] Small humanoids trained in simulation have learned to play soccer, getting up after falls and kicking, and then did it on real robots.[8]
DeXtreme (2022) trained in-hand reorientation policies for an Allegro hand in Isaac Gym, and its vision-based policies beat earlier vision policies on the same task.[7] A 2023 Google DeepMind study trained a 20-joint humanoid to play one-versus-one soccer with deep RL and transferred it zero-shot, crediting high-frequency control, targeted dynamics randomization and training perturbations.[8] In industry, Figure’s Helix 02 relies on System 0, a whole-body controller trained on over 1,000 hours of human motion data plus sim-to-real RL.[9]
Faster simulators and automatic rewards
Simulators have also become much faster. NVIDIA‘s Isaac Gym (2021) put physics simulation and policy training on the same GPU and reported two to three orders of magnitude faster training than CPU-based simulators.[10] Reward functions, the scores that tell a learning robot what counts as success, are normally written by human experts. Eureka (2023) used OpenAI‘s GPT-4 to write reward code instead; across 29 environments and 10 robot types its rewards beat expert-written ones on 83% of tasks.[11]
Simulation and world models in 2026
Simulation is now part of commercial robot platforms. NVIDIA’s Isaac GR00T platform includes the Isaac Sim and Isaac Lab tools alongside its foundation models.[12] In March 2026 NVIDIA announced Isaac Lab 3.0 and Cosmos 3, which it describes as a world foundation model that generates synthetic worlds and simulates actions.[4] GR00T models are trained partly on synthetically generated data.[13]
Limits of simulation
Self-driving offers a caution. Waymo runs closed-loop simulation with its foundation model.[14] It still says “there is simply no substitute” for real-world autonomous experience.[15] For manipulation, the leading generalist models are trained mainly on real robot demonstrations, with simulation and human video as supplements.[16][13]
Questions readers ask
What is the reality gap?
It is the difference between simulated robotics and experiments on real hardware. Sim-to-real methods try to bridge it.[1]
What is domain randomization?
It means randomizing a simulator's settings, such as rendering, during training. With enough variation, the real world looks to the model like just another variation.[2]
Can simulation replace real-world data?
Not entirely. Waymo, which runs a large simulator, says there is no substitute for its volume of real-world autonomous driving experience.[15]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
The "reality gap" is the difference between simulated robotics and experiments on real hardware, which sim-to-real methods try to bridge. confirmedas of 2017-03-20
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World · arXiv · 2017-03-20 · Abstract (retrieved 2026-10-10)
- [2]
Domain randomization, introduced for sim-to-real transfer in 2017, randomizes a simulator's rendering so that with enough variation the real world looks to the model like just another variation. confirmedas of 2017-03-20
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World · arXiv · 2017-03-20 · Abstract (retrieved 2026-10-10)
- [3]
A 2020 Science Robotics study trained a quadruped (ANYmal) locomotion controller by reinforcement learning in simulation and reported zero-shot transfer to natural terrain such as mud, snow and rubble. confirmedas of 2020-10-21
- Learning Quadrupedal Locomotion over Challenging Terrain · arXiv (published in Science Robotics, 2020) · 2020-10-21 · Abstract (retrieved 2026-10-10)
- [4]
In March 2026 NVIDIA announced Cosmos 3, which it describes as a world foundation model combining synthetic world generation, vision reasoning and action simulation, and Isaac Lab 3.0 for large-scale robot learning. confirmedas of 2026-03-16
- NVIDIA and Global Robotics Leaders Take Physical AI to the Real World · NVIDIA Newsroom · 2026-03-16 (retrieved 2026-10-10)
- [5]
In 2019 OpenAI researchers used automatic domain randomization, which generates simulated environments of increasing difficulty, to train a robot hand in simulation that then solved a Rubik's cube on real hardware. confirmedas of 2019-10-16
- Solving Rubik's Cube with a Robot Hand · arXiv (OpenAI authors) · 2019-10-16 · Abstract (retrieved 2026-10-10)
- [6]
A 2023 study trained a transformer-based humanoid walking controller with reinforcement learning in randomized simulated environments and deployed it on a real humanoid outdoors without further fine-tuning. confirmedas of 2023-12-14
- Real-World Humanoid Locomotion with Reinforcement Learning · arXiv · 2023-03-06 · Abstract (retrieved 2026-10-10)
- [7]
DeXtreme (2022) trained in-hand manipulation policies for a multi-fingered Allegro robot hand in Isaac Gym simulation and transferred them to the real hand, with vision-based policies that outperformed earlier vision policies on the same reorientation task. confirmedas of 2022-10-25
- DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality · arXiv · 2022-10-25 (retrieved 2026-10-10)
- [8]
A 2023 Google DeepMind study trained a small 20-joint humanoid to play one-versus-one soccer with deep reinforcement learning in simulation and transferred it to real robots zero-shot, crediting high-frequency control, targeted dynamics randomization and perturbations during training. confirmedas of 2023-04-26
- Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning · arXiv (Google DeepMind authors) · 2023-04-26 (retrieved 2026-10-10)
- [9]
Figure says Helix 02 (January 2026) controls the whole robot from pixels, unloading and reloading a dishwasher across a kitchen in a four-minute autonomous task, using a learned whole-body controller, System 0, trained on more than 1,000 hours of human motion data and sim-to-real reinforcement learning. confirmedas of 2026-01-27
- Introducing Helix 02: Full-Body Autonomy · Figure AI · 2026-01-27 (retrieved 2026-10-10)
- [10]
NVIDIA's Isaac Gym (2021) runs physics simulation and policy training on the GPU, which its authors report speeds up robot reinforcement learning by two to three orders of magnitude compared with CPU-based simulators. confirmedas of 2021-08-24
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning · arXiv (NVIDIA authors) · 2021-08-24 (retrieved 2026-10-10)
- [11]
Eureka (2023) used GPT-4 to write reward functions for reinforcement learning; across 29 environments with 10 robot types, its rewards outperformed expert human-written rewards on 83% of tasks. confirmedas of 2023-10-19
- Eureka: Human-Level Reward Design via Coding Large Language Models · arXiv · 2023-10-19 (retrieved 2026-10-10)
- [12]
NVIDIA's Isaac GR00T platform includes teleoperation tools for capturing demonstrations, open GR00T foundation models, and the Isaac Sim and Isaac Lab simulation and training tools. confirmedas of 2026-05-31
- NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research · NVIDIA Newsroom · 2026-05-31 (retrieved 2026-10-10)
- [13]
GR00T N1 was trained on a mixture of real-robot trajectories, human videos and synthetically generated data, and was deployed on the Fourier GR-1 humanoid. confirmedas of 2025-03-18
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots · arXiv (NVIDIA authors) · 2025-03-18 · Abstract (retrieved 2026-10-10)
- [14]
Waymo describes a Waymo Foundation Model that combines a fast sensor-fusion encoder with a driving vision-language model built on Gemini, and that powers its driver, simulator and critic. confirmedas of 2025-12-09
- Demonstrably Safe AI For Autonomous Driving · Waymo · 2025-12-09 (retrieved 2026-10-10)
- [15]
Waymo says there is no substitute for its volume of real-world fully autonomous driving experience, even with high-fidelity simulation. confirmedas of 2025-12-09
- Demonstrably Safe AI For Autonomous Driving · Waymo · 2025-12-09 (retrieved 2026-10-10)
- [16]
π0 was trained on data from multiple robot types, including single-arm, dual-arm and mobile manipulators, and evaluated on tasks such as laundry folding, table cleaning and box assembly. confirmedas of 2024-10-31
- π0: A Vision-Language-Action Flow Model for General Robot Control · arXiv (Physical Intelligence authors) · 2024-10-31 · Abstract (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Sim-to-real: training robots in simulation for the real world." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/sim-to-real-transfer
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerRobotics and embodied AI in 2026: a crash courseA crash course on robotics and embodied AI as of October 2026: how robots learn, robot foundation models, humanoids, self-driving and the open debates.
- ExplainerHow robots learn: demonstrations, trial and error, and dataA plain guide to how modern robots learn skills from human demonstrations, reinforcement learning and pooled datasets, and why data is the bottleneck.
- WikiNVIDIA Isaac GR00TIsaac GR00T is NVIDIA's platform of open humanoid robot foundation models, simulation tools and reference hardware. Versions, data and partners.
- WikiWaymoWaymo runs driverless robotaxis in US cities. Its scale, safety data and AI approach, and why self-driving cars matter for embodied AI.
- WikiFigure AIFigure AI builds general-purpose humanoid robots and its own Helix AI model. Its robots, BMW deployments, funding and 2026 milestones.
- WikiGemini RoboticsGemini Robotics is Google DeepMind's family of robot AI models. What it does, how it evolved from RT-2 to Gemini Robotics 2, and its limits.