Explainer
How robots learn: demonstrations, trial and error, and data
Most new robot skills are now learned rather than hand-programmed. Robots copy human demonstrations collected by teleoperation (imitation learning)[1] or improve by trial and error in simulation (reinforcement learning)[2]. The limiting factor is data: researchers have pooled about a million real robot trajectories,[3] far less than the text behind language models.[4]
Robots used to be programmed step by step for one job in one fixed cell. The modern approach trains a neural network “policy” that maps what the robot senses to what it does. This page explains the two main ways to train such a policy and why collecting data is still the hardest part.[1][2]
Learning by copying
In imitation learning, a person shows the robot what to do, often by steering it remotely (“teleoperation”), and the robot learns to copy. In a 2023 study a cheap two-armed setup learned six fiddly tasks, such as opening a translucent condiment cup, with 80-90% success after about ten minutes of demonstrations.[1]
A known weakness is that small mistakes add up. Once the robot drifts away from anything it saw in the demonstrations, it may not know how to recover.[5]
Behaviour cloning fits a policy to state-action pairs from expert demonstrations. Its classic failure mode is compounding error under distribution shift. The ACT method (Action Chunking with Transformers) reduces this by predicting chunks of future actions with a generative model.[5] Diffusion Policy treats the policy as a conditional denoising diffusion process. That handles multimodal action distributions, and the paper reported a 46.9% average improvement across 12 tasks from 4 benchmarks.[6] Flow matching, a related generative method, is the action head in π0.[7]
Learning by trial and error
In reinforcement learning (RL), the robot tries actions and gets a reward signal when it does well. Real robots are slow and can break, so much of this training happens in simulation. A four-legged robot trained this way in simulation then walked over mud, snow and rubble in the real world.[2] See the sim-to-real transfer explainer.
RL in randomized simulation has produced well-known legged-locomotion controllers. Examples include the 2020 ANYmal controller[2] and a transformer-based humanoid walking policy that was deployed outdoors without fine-tuning.[8] The leading generalist manipulation models, by contrast, are trained mainly on real robot demonstrations, sometimes mixed with human video and synthetic data.[9][10] RL is now also applied on top of such models in the real world: Physical Intelligence’s Recap method fine-tunes π*0.6 on demonstrations, human corrections and the robot’s own attempts, and the company reports more than doubled throughput on some hard tasks.[11]
The data bottleneck
Language models train on text gathered from the internet. Robots have no equivalent source of action data. Roboticist Ken Goldberg calls the difference a “100,000-year data gap”.[4] One response is to pool data. The Open X-Embodiment collaboration gathered more than a million real trajectories from 22 robot types and 21 institutions.[3] Models trained on the mix showed positive transfer across robots.[12] Newer open datasets add variety: DROID collected 76,000 demonstrations in 564 scenes across three continents.[13]
Companies are also building their own data. NVIDIA trains its GR00T models on real robot data, human videos and synthetic data.[10] Figure reports that pretraining on its human-behaviour dataset raised zero-shot success in new homes from 9% to 56%.[14] In August 2026 Figure turned that into a paid programme, Index, which it says has gathered over 16 million videos from contributors.[15]
From skills to generalists
Pooled data and pretrained vision-language models led to generalist policies: one model for many tasks and, ideally, many robots. These are called vision-language-action models (VLAs).[16][12]
Questions readers ask
What is imitation learning in robotics?
The robot learns a policy by copying human demonstrations, often recorded by teleoperating the robot. One 2023 system learned fine two-handed tasks from about ten minutes of demonstrations.[1]
Why can't robots just learn from the internet like chatbots?
The internet has huge amounts of text but little data on robot movement. Goldberg calls the difference a "100,000-year data gap".[4]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
A 2023 study using a low-cost bimanual teleoperation system and Action Chunking with Transformers (ACT) learned six fine manipulation tasks at 80-90% success from about ten minutes of human demonstrations. confirmedas of 2023-04-23
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware · arXiv · 2023-04-23 · Abstract (retrieved 2026-10-10)
- [2]
A 2020 Science Robotics study trained a quadruped (ANYmal) locomotion controller by reinforcement learning in simulation and reported zero-shot transfer to natural terrain such as mud, snow and rubble. confirmedas of 2020-10-21
- Learning Quadrupedal Locomotion over Challenging Terrain · arXiv (published in Science Robotics, 2020) · 2020-10-21 · Abstract (retrieved 2026-10-10)
- [3]
The Open X-Embodiment dataset pooled more than one million real robot trajectories from 22 robot embodiments, collected through a collaboration of 21 institutions and covering 527 skills. confirmedas of 2025-05-14
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models · arXiv (Open X-Embodiment Collaboration) · 2023-10-13 · Abstract (retrieved 2026-10-10)
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models (project site) · Open X-Embodiment Collaboration (retrieved 2026-10-10)
- [4]
UC Berkeley roboticist Ken Goldberg argued in August 2025 Science Robotics papers that robots face a "100,000-year data gap" compared with the text used to train language models. confirmedas of 2025-08-27
- Are we truly on the verge of the humanoid robot revolution? · UC Berkeley News · 2025-08-27 (retrieved 2026-10-10)
- [5]
The ACT paper identifies compounding errors over time as a key challenge in imitation learning and addresses it by predicting sequences ("chunks") of actions. confirmedas of 2023-04-23
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware · arXiv · 2023-04-23 · Abstract (retrieved 2026-10-10)
- [6]
Diffusion Policy (2023) represents a robot's visuomotor policy as a conditional denoising diffusion process and reported an average 46.9% improvement over prior methods across 12 tasks from 4 benchmarks. confirmedas of 2023-03-07
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion · arXiv · 2023-03-07 · Abstract (retrieved 2026-10-10)
- [7]
Physical Intelligence's π0 (October 2024) is a vision-language-action model that adds a flow-matching action generator on top of a pretrained vision-language model. confirmedas of 2024-10-31
- π0: A Vision-Language-Action Flow Model for General Robot Control · arXiv (Physical Intelligence authors) · 2024-10-31 · Abstract (retrieved 2026-10-10)
- [8]
A 2023 study trained a transformer-based humanoid walking controller with reinforcement learning in randomized simulated environments and deployed it on a real humanoid outdoors without further fine-tuning. confirmedas of 2023-12-14
- Real-World Humanoid Locomotion with Reinforcement Learning · arXiv · 2023-03-06 · Abstract (retrieved 2026-10-10)
- [9]
π0 was trained on data from multiple robot types, including single-arm, dual-arm and mobile manipulators, and evaluated on tasks such as laundry folding, table cleaning and box assembly. confirmedas of 2024-10-31
- π0: A Vision-Language-Action Flow Model for General Robot Control · arXiv (Physical Intelligence authors) · 2024-10-31 · Abstract (retrieved 2026-10-10)
- [10]
GR00T N1 was trained on a mixture of real-robot trajectories, human videos and synthetically generated data, and was deployed on the Fourier GR-1 humanoid. confirmedas of 2025-03-18
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots · arXiv (NVIDIA authors) · 2025-03-18 · Abstract (retrieved 2026-10-10)
- [11]
In November 2025 Physical Intelligence introduced π*0.6, trained with a reinforcement-learning method called Recap that combines demonstrations, corrections and autonomous experience; the company says it more than doubled throughput on some of the hardest tasks. confirmedas of 2025-11-17
- π*0.6: a VLA that Learns from Experience · Physical Intelligence · 2025-11-17 (retrieved 2026-10-10)
- [12]
The RT-X models trained on Open X-Embodiment (paper first posted October 2023) showed positive transfer, improving the capabilities of multiple robots by using experience from other robot platforms. confirmedas of 2023-10-13
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models · arXiv (Open X-Embodiment Collaboration) · 2023-10-13 · Abstract (retrieved 2026-10-10)
- [13]
The DROID dataset (2024) contains 76,000 demonstration trajectories, or 350 hours of interaction, collected across 564 scenes and 84 tasks by 50 collectors in North America, Asia and Europe over 12 months, and is fully open source. confirmedas of 2024-03-19
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset · arXiv · 2024-03-19 (retrieved 2026-10-10)
- [14]
On September 17, 2026 Figure reported that Helix 2.5 performed tidying, towel folding and bed making in 30 Bay Area homes where no data had been collected, with 56% zero-shot success versus 9% without pretraining on its Index human-behavior dataset. confirmedas of 2026-09-17
- Helix 2.5: Zero-Shot 30-Home Generalization · Figure AI · 2026-09-17 (retrieved 2026-10-10)
- [15]
In August 2026 Figure launched Index, an app that pays people to record videos for robot training; Figure said it had paid $15 million to contributors, received over 16 million videos and committed to spend over $1 billion on data and compute in the next 12 months. confirmedas of 2026-08-25
- Introducing Index: Building The World's Largest and Most Diverse Physical Dataset · Figure AI · 2026-08-25 (retrieved 2026-10-10)
- [16]
Google DeepMind's RT-2 (July 2023) trained a single vision-language model to output robot actions by expressing those actions as text tokens. confirmedas of 2023-07-28
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · arXiv (Google DeepMind authors) · 2023-07-28 · Abstract (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Added reinforcement learning on real robots (π*0.6), the DROID dataset and Figure's Index data programme.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"How robots learn: demonstrations, trial and error, and data." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-robots-learn
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerRobotics and embodied AI in 2026: a crash courseA crash course on robotics and embodied AI as of October 2026: how robots learn, robot foundation models, humanoids, self-driving and the open debates.
- ExplainerVision-language-action models: how robot foundation models workWhat vision-language-action (VLA) models are, how they turn camera images and instructions into robot motion, and the main models as of 2026.
- ExplainerSim-to-real: training robots in simulation for the real worldHow robots are trained in simulated worlds and then moved to real hardware, what the "reality gap" is, and how domain randomization bridges it.
- WikiOpen X-EmbodimentOpen X-Embodiment is a pooled dataset of 1M+ robot trajectories from 22 robot types and the RT-X models trained on it. Why it mattered.
- WikiFigure AIFigure AI builds general-purpose humanoid robots and its own Helix AI model. Its robots, BMW deployments, funding and 2026 milestones.
- WikiGemini RoboticsGemini Robotics is Google DeepMind's family of robot AI models. What it does, how it evolved from RT-2 to Gemini Robotics 2, and its limits.