Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    Sim-to-real: training robots in simulation for the real world

    Sim-to-real transfer means training a robot controller in simulation and deploying it on real hardware. The difference between the two is called the "reality gap".[1] Randomizing the simulator helps bridge it, because a model that has seen enough variation treats the real world as one more variation.[2] The method trained walking robots that worked outdoors with no real-world training,[3] and NVIDIA and others now sell simulation and "world model" tools for robot learning.[4]

    Editor reviewedUpdated Robotics and embodied AIArtificial intelligence

    Training a robot in the physical world is slow and expensive. Simulation is fast and safe, but never quite matches reality. Sim-to-real research is about closing that “reality gap”.[1] This page explains the main technique and where it works best.

    The reality gap

    A simulator is a video-game-like world with physics. A robot can practise millions of times there without breaking anything. But simulated friction, lighting and motors are never exactly like real ones. A controller that works perfectly in simulation can fail on a real robot. Researchers call this mismatch the “reality gap”.[1]

    Domain randomization, introduced for vision in 2017, randomizes rendering so that “the real world may appear to the model as just another variation”.[2] OpenAI’s automatic domain randomization grew the randomization range as the policy improved. It trained a robot hand in simulation that then solved a Rubik’s cube on real hardware.[5]

    Where it works: walking

    Sim-to-real works especially well for walking. A four-legged robot trained only in simulation walked over mud, snow and rubble outdoors without any extra training.[3] A human-sized humanoid was later taught to walk the same way.[6]

    The 2020 ANYmal controller was trained with RL in simulation and showed zero-shot generalization to natural terrain.[3] A 2023 humanoid controller used a causal transformer over the history of proprioceptive observations and actions, trained with RL across randomized simulated environments. It adapted its behaviour in context on real terrain without weight updates.[6]

    Where it works: hands and whole bodies

    Simulation also works for hands. Researchers trained a robot hand in simulation to turn objects in its fingers and then used it on a real hand.[7] Small humanoids trained in simulation have learned to play soccer, getting up after falls and kicking, and then did it on real robots.[8]

    DeXtreme (2022) trained in-hand reorientation policies for an Allegro hand in Isaac Gym, and its vision-based policies beat earlier vision policies on the same task.[7] A 2023 Google DeepMind study trained a 20-joint humanoid to play one-versus-one soccer with deep RL and transferred it zero-shot, crediting high-frequency control, targeted dynamics randomization and training perturbations.[8] In industry, Figure’s Helix 02 relies on System 0, a whole-body controller trained on over 1,000 hours of human motion data plus sim-to-real RL.[9]

    Faster simulators and automatic rewards

    Simulators have also become much faster. NVIDIA‘s Isaac Gym (2021) put physics simulation and policy training on the same GPU and reported two to three orders of magnitude faster training than CPU-based simulators.[10] Reward functions, the scores that tell a learning robot what counts as success, are normally written by human experts. Eureka (2023) used OpenAI‘s GPT-4 to write reward code instead; across 29 environments and 10 robot types its rewards beat expert-written ones on 83% of tasks.[11]

    Simulation and world models in 2026

    Simulation is now part of commercial robot platforms. NVIDIA’s Isaac GR00T platform includes the Isaac Sim and Isaac Lab tools alongside its foundation models.[12] In March 2026 NVIDIA announced Isaac Lab 3.0 and Cosmos 3, which it describes as a world foundation model that generates synthetic worlds and simulates actions.[4] GR00T models are trained partly on synthetically generated data.[13]

    Limits of simulation

    Self-driving offers a caution. Waymo runs closed-loop simulation with its foundation model.[14] It still says “there is simply no substitute” for real-world autonomous experience.[15] For manipulation, the leading generalist models are trained mainly on real robot demonstrations, with simulation and human video as supplements.[16][13]

    Questions readers ask

    What is the reality gap?

    It is the difference between simulated robotics and experiments on real hardware. Sim-to-real methods try to bridge it.[1]

    What is domain randomization?

    It means randomizing a simulator's settings, such as rendering, during training. With enough variation, the real world looks to the model like just another variation.[2]

    Can simulation replace real-world data?

    Not entirely. Waymo, which runs a large simulator, says there is no substitute for its volume of real-world autonomous driving experience.[15]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      The "reality gap" is the difference between simulated robotics and experiments on real hardware, which sim-to-real methods try to bridge. confirmedas of 2017-03-20

    2. [2]

      Domain randomization, introduced for sim-to-real transfer in 2017, randomizes a simulator's rendering so that with enough variation the real world looks to the model like just another variation. confirmedas of 2017-03-20

    3. [3]

      A 2020 Science Robotics study trained a quadruped (ANYmal) locomotion controller by reinforcement learning in simulation and reported zero-shot transfer to natural terrain such as mud, snow and rubble. confirmedas of 2020-10-21

    4. [4]

      In March 2026 NVIDIA announced Cosmos 3, which it describes as a world foundation model combining synthetic world generation, vision reasoning and action simulation, and Isaac Lab 3.0 for large-scale robot learning. confirmedas of 2026-03-16

    5. [5]

      In 2019 OpenAI researchers used automatic domain randomization, which generates simulated environments of increasing difficulty, to train a robot hand in simulation that then solved a Rubik's cube on real hardware. confirmedas of 2019-10-16

    6. [6]

      A 2023 study trained a transformer-based humanoid walking controller with reinforcement learning in randomized simulated environments and deployed it on a real humanoid outdoors without further fine-tuning. confirmedas of 2023-12-14

    7. [7]

      DeXtreme (2022) trained in-hand manipulation policies for a multi-fingered Allegro robot hand in Isaac Gym simulation and transferred them to the real hand, with vision-based policies that outperformed earlier vision policies on the same reorientation task. confirmedas of 2022-10-25

    8. [8]

      A 2023 Google DeepMind study trained a small 20-joint humanoid to play one-versus-one soccer with deep reinforcement learning in simulation and transferred it to real robots zero-shot, crediting high-frequency control, targeted dynamics randomization and perturbations during training. confirmedas of 2023-04-26

    9. [9]

      Figure says Helix 02 (January 2026) controls the whole robot from pixels, unloading and reloading a dishwasher across a kitchen in a four-minute autonomous task, using a learned whole-body controller, System 0, trained on more than 1,000 hours of human motion data and sim-to-real reinforcement learning. confirmedas of 2026-01-27

    10. [10]

      NVIDIA's Isaac Gym (2021) runs physics simulation and policy training on the GPU, which its authors report speeds up robot reinforcement learning by two to three orders of magnitude compared with CPU-based simulators. confirmedas of 2021-08-24

    11. [11]

      Eureka (2023) used GPT-4 to write reward functions for reinforcement learning; across 29 environments with 10 robot types, its rewards outperformed expert human-written rewards on 83% of tasks. confirmedas of 2023-10-19

    12. [12]

      NVIDIA's Isaac GR00T platform includes teleoperation tools for capturing demonstrations, open GR00T foundation models, and the Isaac Sim and Isaac Lab simulation and training tools. confirmedas of 2026-05-31

    13. [13]

      GR00T N1 was trained on a mixture of real-robot trajectories, human videos and synthetically generated data, and was deployed on the Fourier GR-1 humanoid. confirmedas of 2025-03-18

    14. [14]

      Waymo describes a Waymo Foundation Model that combines a fast sensor-fusion encoder with a driving vision-language model built on Gemini, and that powers its driver, simulator and critic. confirmedas of 2025-12-09

    15. [15]

      Waymo says there is no substitute for its volume of real-world fully autonomous driving experience, even with high-fidelity simulation. confirmedas of 2025-12-09

    16. [16]

      π0 was trained on data from multiple robot types, including single-arm, dual-arm and mobile manipulators, and evaluated on tasks such as laundry folding, table cleaning and box assembly. confirmedas of 2024-10-31

    Revision history (2)
    1. Page created.
    2. Added GPU simulation (Isaac Gym), dexterous hands (DeXtreme), humanoid soccer, LLM-written rewards (Eureka) and Figure's sim-to-real whole-body controller.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Sim-to-real: training robots in simulation for the real world." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/sim-to-real-transfer

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.