Explainer
Robotics and embodied AI in 2026: a crash course
Embodied AI is AI that senses and acts in the physical world through robots and self-driving cars. Since 2023, robotics has borrowed the foundation-model recipe from language AI: vision-language-action (VLA) models turn camera images and instructions into robot motion.[1] By October 2026, such models control whole humanoid bodies,[2] humanoids do narrow factory jobs,[3] and robotaxis give hundreds of thousands of paid rides a week.[4] Reliability, data and demand remain open questions.[5][6]
Robotics has adopted the recipe behind modern language AI. Instead of hand-coding each skill, researchers train large general models on broad data and adapt them to many tasks and robot bodies.[7][1] This page is the map. The course list takes you from the basics to the October 2026 frontier.
Why it matters
Much everyday work is physical: moving, assembling, cleaning, driving. Software AI cannot do these jobs; robots can. If general-purpose robots become reliable, they could change factories, warehouses and homes. That is why money is flowing in. Figure raised more than $1 billion at a $39 billion valuation in 2025,[8] and China’s Unitree raised about $900 million in an August 2026 IPO.[9] Governments are paying attention too. On 28 July 2026 the US FCC added foreign-made humanoids and other mobile robots to its Covered List on security grounds, blocking approval of new models.[10]
The thesis behind the investment is that scaling laws carry over from language to action. Waymo has reported power-law scaling in driving models,[11] and pooled multi-robot data shows positive transfer.[7] The counter-thesis is that robot data is scarce and costly, and that industrial reliability (around 99.99%) is far above current VLA success rates of roughly 45-76% on hard tasks.[6][5][12]
The map
The field has four overlapping parts:
- Robot learning. How robots acquire skills: imitation of human demonstrations, reinforcement learning, and pooled datasets such as Open X-Embodiment.[13][14]
- Simulation and sim-to-real. Training in virtual worlds, then transferring to hardware across the “reality gap”.[15] NVIDIA now offers world models such as Cosmos 3 for this.[16]
- Robot foundation models (VLAs). Large models that map images and language to actions, such as Gemini Robotics, GR00T, Physical Intelligence‘s π models and Figure’s Helix.[2][17][18][19]
- Embodiments. The physical platforms: humanoids, arms, quadrupeds and self-driving cars such as Waymo‘s robotaxis.[20][4]
Key ideas
Policy: the program that decides what the robot does next, based on what it senses. Imitation learning: the robot copies human demonstrations.[13] Reinforcement learning: the robot learns by trial and error, usually in simulation.[21] VLA: a vision-language model that outputs actions instead of words.[1] Cross-embodiment: one model learns from many robot types.[7]
Action representation has moved from discrete tokens (RT-2) to continuous generative heads: diffusion (Diffusion Policy, GR00T N1) and flow matching (π0).[1][22][23][18] A common design pairs a slow vision-language “System 2” with a fast motor “System 1”. Helix, for example, runs at 7-9 Hz and 200 Hz respectively.[19] Data strategies mix teleoperation, human video and synthetic data.[24][25]
Who the main players are
- Model builders. Google DeepMind (Gemini Robotics 2, July 2026),[2] NVIDIA (Isaac GR00T, with N1.7 openly licensed and N2 due by the end of 2026)[17][26] and Physical Intelligence, whose π0.7 arrived in April 2026.[27][28]
- Humanoid makers. Figure (deployed at BMW),[3] Tesla Optimus (factory under construction),[29] and Chinese volume leaders Unitree and AGIBOT.[20] Others include Agility, Apptronik, Boston Dynamics and 1X.[30][31]
- Self-driving. Waymo, with more than 270 million fully autonomous miles.[32]
- Academia and open source. Shared datasets and open models, such as Open X-Embodiment (21 institutions), OpenVLA and DROID, plus Hugging Face’s small SmolVLA.[14][33][34][35]
The frontier in October 2026
Three results mark the frontier. Gemini Robotics 2 controls a full humanoid, walking and manipulating from one instruction.[36] Figure’s Helix 2.5 did household tasks in 30 unseen homes, with 56% zero-shot success.[25] NVIDIA claims GR00T N2 doubles success on new tasks in new environments.[17] On deployment, Figure 03 works in BMW logistics,[3] Tesla is still installing its Optimus line,[29] and China leads shipments, with about 85% of the market.[37]
The main open questions are whether data scaling will close the reliability gap, whether there is enough demand for humanoids, and how safety standards and trade policy will shape the market. The debate page covers them, and the tracker follows new milestones.[38][39][10]
How to use this course
Read in order: fundamentals first, then the technology and company pages, the debate and the tracker. Company figures are usually self-reported.[25]
Questions readers ask
What is embodied AI?
AI that acts in the physical world through a body, such as a robot arm, a humanoid or a self-driving car. Modern systems often use vision-language-action models that turn images and instructions into motor commands.[1][11]
Who are the main players in 2026?
In robot AI models, Google DeepMind (Gemini Robotics), NVIDIA (Isaac GR00T) and Physical Intelligence. In humanoids, Figure, Tesla and China's Unitree and AGIBOT. In self-driving, Waymo.[2][17][40][20][4]
Are humanoid robots being used for real work?
In narrow roles. Figure 03 began parts sequencing at BMW's Spartanburg plant in June 2026. About 15,000 humanoids shipped worldwide in 2025.[3][20]
Why is robotics behind chatbots?
Robots lack internet-scale training data. Ken Goldberg calls the shortfall a "100,000-year data gap" compared with the text used to train language models.[6]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
Google DeepMind's RT-2 (July 2023) trained a single vision-language model to output robot actions by expressing those actions as text tokens. confirmedas of 2023-07-28
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · arXiv (Google DeepMind authors) · 2023-07-28 · Abstract (retrieved 2026-10-10)
- [2]
On July 30, 2026 Google DeepMind announced Gemini Robotics 2, a vision-language-action model it says can control full humanoids, including walking and manipulation, as well as bi-arm robots; previous models in the family controlled a humanoid's upper body for tabletop tasks. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 (retrieved 2026-10-10)
- [3]
In June 2026 BMW and Figure began using Figure 03 for parts sequencing in logistics at the Spartanburg plant, picking unsorted components into trolleys for assembly workers. confirmedas of 2026-06-30
- BMW Group advances the use of Physical AI in production with Figure 03 project in Spartanburg · BMW Group PressClub · 2026-06-25 (retrieved 2026-10-10)
- F.03 Arrives at BMW · Figure AI · 2026-06-30 (retrieved 2026-10-10)
- [4]
As of March 2026 Waymo was giving 500,000 paid robotaxi rides a week across 10 US cities, up from about 50,000 a week in May 2024. confirmedas of 2026-03-27
- Waymo's skyrocketing ridership in one chart · TechCrunch · 2026-03-27 (retrieved 2026-10-10)
- [5]
IEEE Spectrum reported that industrial customers expect about 99.99% reliability and that ISO safety standards for dynamically balancing legged robots were still being developed in 2025. reportedas of 2025-09-11
- Reality Is Ruining the Humanoid Robot Hype · IEEE Spectrum · 2025-09-11 (retrieved 2026-10-10)
- [6]
UC Berkeley roboticist Ken Goldberg argued in August 2025 Science Robotics papers that robots face a "100,000-year data gap" compared with the text used to train language models. confirmedas of 2025-08-27
- Are we truly on the verge of the humanoid robot revolution? · UC Berkeley News · 2025-08-27 (retrieved 2026-10-10)
- [7]
The RT-X models trained on Open X-Embodiment (paper first posted October 2023) showed positive transfer, improving the capabilities of multiple robots by using experience from other robot platforms. confirmedas of 2023-10-13
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models · arXiv (Open X-Embodiment Collaboration) · 2023-10-13 · Abstract (retrieved 2026-10-10)
- [8]
In September 2025 Figure said it had raised more than $1 billion in Series C funding at a $39 billion post-money valuation, in a round led by Parkway Venture Capital. confirmedas of 2025-09-16
- Figure Exceeds $1B in Series C Funding at $39B Post-Money Valuation · Figure AI · 2025-09-16 (retrieved 2026-10-10)
- [9]
Unitree's shares rose more than 460% on their first trading day on August 19, 2026, after an IPO that raised about $900 million at an initial valuation of about $9 billion. confirmedas of 2026-08-19
- Unitree, famous for its dancing robots, surges by 460% on its trading debut · Fortune · 2026-08-19 (retrieved 2026-10-10)
- [10]
On 28 July 2026 the FCC added foreign-produced "advanced robotic devices", defined as mobile robots such as humanoids and quadrupeds, to its Covered List after an interagency national-security determination; the restriction applies to new device models, not to devices already bought or models previously approved. confirmedas of 2026-07-28
- FACT SHEET: FCC Updates Covered List to Include Foreign-Produced Advanced Robotic Devices and Power Inverters · Federal Communications Commission · 2026-07-28 (retrieved 2026-10-10)
- FACT SHEET: FCC Updates Covered List to Include Foreign-Produced Advanced Robotic Devices and Power Inverters · Federal Communications Commission · 2026-07-28 (retrieved 2026-10-10)
- U.S. bans foreign-made humanoid robots, targeting China over national security · PBS NewsHour (Associated Press) · 2026-07-29 (retrieved 2026-10-10)
- [11]
Waymo reported in June 2025 that motion-forecasting quality follows a power law in training compute and that closed-loop driving performance follows a similar trend, using a 500,000-hour driving dataset, and said the findings have applications in embodied AI research, including robotics. confirmedas of 2025-06-13
- New Insights for Scaling Laws in Autonomous Driving · Waymo · 2025-06-13 (retrieved 2026-10-10)
- [12]
Google DeepMind reported Gemini Robotics 2 success rates of 45.7% to 76.3% on general whole-body manipulation evaluations and 32% to 92% on multi-finger dexterity tasks. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 · Results section (retrieved 2026-10-10)
- [13]
A 2023 study using a low-cost bimanual teleoperation system and Action Chunking with Transformers (ACT) learned six fine manipulation tasks at 80-90% success from about ten minutes of human demonstrations. confirmedas of 2023-04-23
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware · arXiv · 2023-04-23 · Abstract (retrieved 2026-10-10)
- [14]
The Open X-Embodiment dataset pooled more than one million real robot trajectories from 22 robot embodiments, collected through a collaboration of 21 institutions and covering 527 skills. confirmedas of 2025-05-14
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models · arXiv (Open X-Embodiment Collaboration) · 2023-10-13 · Abstract (retrieved 2026-10-10)
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models (project site) · Open X-Embodiment Collaboration (retrieved 2026-10-10)
- [15]
The "reality gap" is the difference between simulated robotics and experiments on real hardware, which sim-to-real methods try to bridge. confirmedas of 2017-03-20
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World · arXiv · 2017-03-20 · Abstract (retrieved 2026-10-10)
- [16]
In March 2026 NVIDIA announced Cosmos 3, which it describes as a world foundation model combining synthetic world generation, vision reasoning and action simulation, and Isaac Lab 3.0 for large-scale robot learning. confirmedas of 2026-03-16
- NVIDIA and Global Robotics Leaders Take Physical AI to the Real World · NVIDIA Newsroom · 2026-03-16 (retrieved 2026-10-10)
- [17]
On March 16, 2026 NVIDIA made GR00T N1.7 available in early access with commercial licensing and previewed GR00T N2, which it says succeeds at new tasks in new environments more than twice as often as leading VLA models and is slated for release by the end of 2026. confirmedas of 2026-03-16
- NVIDIA and Global Robotics Leaders Take Physical AI to the Real World · NVIDIA Newsroom · 2026-03-16 (retrieved 2026-10-10)
- [18]
Physical Intelligence's π0 (October 2024) is a vision-language-action model that adds a flow-matching action generator on top of a pretrained vision-language model. confirmedas of 2024-10-31
- π0: A Vision-Language-Action Flow Model for General Robot Control · arXiv (Physical Intelligence authors) · 2024-10-31 · Abstract (retrieved 2026-10-10)
- [19]
Figure's Helix (February 2025) is a vision-language-action model with a slower onboard vision-language "System 2" at 7-9 Hz and a fast visuomotor "System 1" that outputs actions at 200 Hz. confirmedas of 2025-02-20
- Helix: A Vision-Language-Action Model for Generalist Humanoid Control · Figure AI · 2025-02-20 (retrieved 2026-10-10)
- [20]
About 15,000 humanoid robots were shipped worldwide in 2025; Unitree and AGIBOT each shipped more than 5,000, while Tesla and Figure each shipped a few hundred or fewer. reportedas of 2026-07-29
- U.S. bans foreign-made humanoid robots, targeting China over national security · PBS NewsHour (Associated Press) · 2026-07-29 (retrieved 2026-10-10)
- [21]
A 2020 Science Robotics study trained a quadruped (ANYmal) locomotion controller by reinforcement learning in simulation and reported zero-shot transfer to natural terrain such as mud, snow and rubble. confirmedas of 2020-10-21
- Learning Quadrupedal Locomotion over Challenging Terrain · arXiv (published in Science Robotics, 2020) · 2020-10-21 · Abstract (retrieved 2026-10-10)
- [22]
Diffusion Policy (2023) represents a robot's visuomotor policy as a conditional denoising diffusion process and reported an average 46.9% improvement over prior methods across 12 tasks from 4 benchmarks. confirmedas of 2023-03-07
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion · arXiv · 2023-03-07 · Abstract (retrieved 2026-10-10)
- [23]
GR00T N1 (March 2025) is an open humanoid foundation model with a vision-language module for understanding and a diffusion-transformer module that generates motor actions in real time. confirmedas of 2025-03-18
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots · arXiv (NVIDIA authors) · 2025-03-18 · Abstract (retrieved 2026-10-10)
- [24]
GR00T N1 was trained on a mixture of real-robot trajectories, human videos and synthetically generated data, and was deployed on the Fourier GR-1 humanoid. confirmedas of 2025-03-18
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots · arXiv (NVIDIA authors) · 2025-03-18 · Abstract (retrieved 2026-10-10)
- [25]
On September 17, 2026 Figure reported that Helix 2.5 performed tidying, towel folding and bed making in 30 Bay Area homes where no data had been collected, with 56% zero-shot success versus 9% without pretraining on its Index human-behavior dataset. confirmedas of 2026-09-17
- Helix 2.5: Zero-Shot 30-Home Generalization · Figure AI · 2026-09-17 (retrieved 2026-10-10)
- [26]
As of October 2026 GR00T N1.7 was generally available under the Apache 2.0 licence, with a new vision-language backbone and 20,000 hours of human video in pretraining; NVIDIA says it performs comparably to N1.6 with better generalization and language following. confirmedas of 2026-10-10
- NVIDIA/Isaac-GR00T: NVIDIA Isaac GR00T N1.7 (README) · NVIDIA (GitHub) (retrieved 2026-10-10)
- [27]
Physical Intelligence's π0.5 (April 2025) uses co-training on heterogeneous tasks and multimodal examples (images, language commands, object detections, semantic subtask predictions and low-level actions) to perform long-horizon tasks such as cleaning a kitchen or bedroom in homes not seen during training. confirmedas of 2025-04-22
- π0.5: a Vision-Language-Action Model with Open-World Generalization · arXiv (Physical Intelligence authors) · 2025-04-22 · Abstract (retrieved 2026-10-10)
- [28]
Physical Intelligence said in April 2026 that its π0.7 model performs dexterous tasks as well as fine-tuned specialists and shows first signs of compositional generalization, such as folding laundry on a robot with no laundry-folding data. confirmedas of 2026-04-16
- π0.7: a Steerable Model with Emergent Capabilities · Physical Intelligence · 2026-04-16 (retrieved 2026-10-10)
- [29]
In its Q2 2026 update (July 2026) Tesla said it had decommissioned the Model S and X lines, was installing first-generation Optimus lines with production anticipated later in 2026, and that initial builds would be used for training-data collection. confirmedas of 2026-07-22
- Tesla Q2 2026 Update (Form 8-K, Exhibit 99.1) · Tesla (SEC filing) · 2026-07-22 (retrieved 2026-10-10)
- [30]
NVIDIA named 1X, AGIBOT, Agility, Boston Dynamics, Figure and others among humanoid developers building on its robotics platform in March 2026. confirmedas of 2026-03-16
- NVIDIA and Global Robotics Leaders Take Physical AI to the Real World · NVIDIA Newsroom · 2026-03-16 (retrieved 2026-10-10)
- [31]
Google DeepMind lists Agile Robots, Apptronik and Boston Dynamics among its Gemini Robotics partners and runs a program with more than 100 trusted testers. confirmedas of 2026-10-10
- Gemini Robotics (model page) · Google DeepMind (retrieved 2026-10-10)
- [32]
Waymo reports more than 270 million fully autonomous miles through June 2026, with 82% fewer injury-causing crashes and 95% fewer serious-injury-or-worse crashes than human benchmarks in its main cities. confirmedas of 2026-09-24
- September 24, 2026 - From the road: safety data update · Waymo · 2026-09-24 (retrieved 2026-10-10)
- [33]
OpenVLA (2024) is an open-source 7-billion-parameter vision-language-action model trained on 970,000 real-world robot demonstrations from Open X-Embodiment. confirmedas of 2024-06-13
- OpenVLA: An Open-Source Vision-Language-Action Model · arXiv · 2024-06-13 · Abstract (retrieved 2026-10-10)
- [34]
The DROID dataset (2024) contains 76,000 demonstration trajectories, or 350 hours of interaction, collected across 564 scenes and 84 tasks by 50 collectors in North America, Asia and Europe over 12 months, and is fully open source. confirmedas of 2024-03-19
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset · arXiv · 2024-03-19 (retrieved 2026-10-10)
- [35]
Hugging Face's SmolVLA (June 2025) is a small VLA trained on community-collected data that can be trained on a single GPU and deployed on consumer GPUs or CPUs, and its authors report performance comparable to VLAs ten times larger. confirmedas of 2025-06-02
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics · arXiv (Hugging Face authors) · 2025-06-02 (retrieved 2026-10-10)
- [36]
Google DeepMind demonstrated Gemini Robotics 2 on Apptronik's Apollo 2 humanoid, which walked to a table, picked up a watering can and placed it on a shelf from a spoken instruction. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 (retrieved 2026-10-10)
- [38]
IEEE Spectrum reported in September 2025 that a former Agility Robotics product chief saw demand, not manufacturing, as the main obstacle, saying no one had found an application needing thousands of humanoids per facility. reportedas of 2025-09-11
- Reality Is Ruining the Humanoid Robot Hype · IEEE Spectrum · 2025-09-11 (retrieved 2026-10-10)
- [39]
Goldberg advocates combining conventional engineering with learning, so that robots doing useful work (as Waymo and Ambi Robotics do) collect data that improves them over time. confirmedas of 2025-08-27
- Are we truly on the verge of the humanoid robot revolution? · UC Berkeley News · 2025-08-27 (retrieved 2026-10-10)
- [40]
In March 2026 Bloomberg News reported that Physical Intelligence was in talks to raise about $1 billion at a valuation above $11 billion, roughly double its previous $5.6 billion valuation. reportedas of 2026-03-29
- Physical Intelligence Seeks $1 Billion as Robotics Interest Grows · PYMNTS (citing Bloomberg News) · 2026-03-29 (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Corrected the FCC action to the 28 July 2026 Covered List update, citing FCC documents.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Robotics and embodied AI in 2026: a crash course." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/robotics
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- DevelopingRobotics and embodied AI tracker: milestones to watchA running timeline of milestones in robot foundation models, humanoids and self-driving since the 2023 start of the VLA era, updated as of October 2026.
- AnalysisAre humanoid robots ready for real work? The 2026 debateEvidence on whether humanoid robots and robot foundation models are ready for factories and homes, with competing views on data, demand and safety.
- WikiFigure AIFigure AI builds general-purpose humanoid robots and its own Helix AI model. Its robots, BMW deployments, funding and 2026 milestones.
- WikiGemini RoboticsGemini Robotics is Google DeepMind's family of robot AI models. What it does, how it evolved from RT-2 to Gemini Robotics 2, and its limits.
- WikiNVIDIA Isaac GR00TIsaac GR00T is NVIDIA's platform of open humanoid robot foundation models, simulation tools and reference hardware. Versions, data and partners.
- WikiOpen X-EmbodimentOpen X-Embodiment is a pooled dataset of 1M+ robot trajectories from 22 robot types and the RT-X models trained on it. Why it mattered.