Explainer
Vision-language-action models: how robot foundation models work
A vision-language-action (VLA) model is a neural network that takes camera images and a language instruction and outputs robot actions. It is usually built on a pretrained vision-language model.[1][2] Since RT-2 in 2023, VLAs have become the main approach to general-purpose robot control. As of 2026, models from Google DeepMind, NVIDIA, Physical Intelligence and Figure control whole humanoid bodies.[3][4]
Vision-language-action models are the robotics version of the “foundation model” idea. One large network, pretrained on broad data, is adapted to control robots across many tasks.[1][5] They are the core technology behind most 2025-2026 humanoid demonstrations.[3][4]
The basic idea
A chatbot reads words and writes words. A VLA looks through the robot’s cameras, reads an instruction such as “put the watering can in the green bin”, and writes out motor commands instead of text.[6] It starts from a model that already understands images and language from web data. That helps it handle objects and commands it never saw during robot training.[7]
Google DeepMind’s RT-2 (2023) co-fine-tuned a vision-language model on web and robot data by expressing robot actions as text tokens. The resulting model generalized to novel objects and showed rudimentary semantic reasoning.[1][7] OpenVLA (2024) is an open 7-billion-parameter version, trained on 970,000 demonstrations from Open X-Embodiment. It reported 16.5% higher absolute success than the 55-billion-parameter RT-2-X across 29 tasks.[8][9]
Fast hands, slow thinking
Robots must move smoothly many times a second, but big language models are slow. Many VLAs therefore split the work in two. A slower “thinking” part understands the scene and the instruction, and a fast “moving” part turns that into motion.[10][11]
Discrete action tokens limit control frequency and precision, so newer models attach continuous action heads. π0 uses flow matching on top of a pretrained VLM.[2] NVIDIA’s GR00T N1 pairs a vision-language module with a diffusion transformer that generates actions in real time.[11] Figure’s Helix runs an onboard VLM (“System 2”) at 7-9 Hz that conditions a visuomotor policy (“System 1”) outputting actions at 200 Hz.[10] Google DeepMind separates an embodied-reasoning model (Gemini Robotics-ER) from the action model.[12]
Main models as of October 2026
- Gemini Robotics (gemini-robotics): Google DeepMind‘s family. Gemini Robotics 1.5 (September 2025) “thinks” before acting, an On-Device version runs locally on bi-arm robots, and Gemini Robotics 2 (July 2026) is the first in the family to control whole humanoids, including walking.[13][14][3]
- π series (physical-intelligence): π0 (2024) and π0.5 (2025), which cleaned kitchens and bedrooms in homes not seen during training.[2][15] π*0.6 (November 2025) added reinforcement learning from the robot’s own experience, and π0.7 (April 2026) showed early signs of combining skills in new ways.[16][17]
- GR00T (isaac-gr00t): NVIDIA‘s open humanoid models. N1.6 (January 2026) added full-body control; N1.7 is generally available under Apache 2.0; N2 is due by the end of 2026.[18][19][20]
- Helix (figure-ai): Figure’s in-house VLA. Helix 02 (January 2026) controls the whole body, and Helix 2.5 was tested zero-shot in 30 homes.[21][22]
- Open research models: OpenVLA (7 billion parameters) and Hugging Face’s SmolVLA, which runs on consumer GPUs or even CPUs.[8][23]
Turning motions into tokens
Language models predict tokens, so a VLA needs a way to express continuous motions as tokens or to generate them directly. RT-2 wrote actions as text tokens.[1] Physical Intelligence’s FAST tokenizer compresses action sequences with the discrete cosine transform and, with π0, matched diffusion-based models while training up to five times faster.[24] Others skip tokens: GR00T uses a diffusion-transformer head that denoises continuous actions.[25]
Limits
Reported success rates show how far these systems still are from industrial reliability. Gemini Robotics 2 scored 45.7% to 76.3% on general whole-body manipulation evaluations,[26] and Figure’s Helix 2.5 reached 56% zero-shot success in unseen homes.[22] Industrial customers are reported to expect around 99.99% reliability.[27] These figures come from the developers’ own evaluations, so they are not directly comparable across companies.[26][22]
Questions readers ask
What does "vision-language-action" mean?
The model takes images (vision) and an instruction (language) as input and outputs robot motor commands (action). RT-2 did this by encoding the actions as text tokens.[1]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
Google DeepMind's RT-2 (July 2023) trained a single vision-language model to output robot actions by expressing those actions as text tokens. confirmedas of 2023-07-28
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · arXiv (Google DeepMind authors) · 2023-07-28 · Abstract (retrieved 2026-10-10)
- [2]
Physical Intelligence's π0 (October 2024) is a vision-language-action model that adds a flow-matching action generator on top of a pretrained vision-language model. confirmedas of 2024-10-31
- π0: A Vision-Language-Action Flow Model for General Robot Control · arXiv (Physical Intelligence authors) · 2024-10-31 · Abstract (retrieved 2026-10-10)
- [3]
On July 30, 2026 Google DeepMind announced Gemini Robotics 2, a vision-language-action model it says can control full humanoids, including walking and manipulation, as well as bi-arm robots; previous models in the family controlled a humanoid's upper body for tabletop tasks. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 (retrieved 2026-10-10)
- [4]
Figure says Helix 02 coordinates Figure 03's hands, arms, torso and feet. confirmedas of 2026-06-30
- F.03 Arrives at BMW · Figure AI · 2026-06-30 (retrieved 2026-10-10)
- [5]
The RT-X models trained on Open X-Embodiment (paper first posted October 2023) showed positive transfer, improving the capabilities of multiple robots by using experience from other robot platforms. confirmedas of 2023-10-13
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models · arXiv (Open X-Embodiment Collaboration) · 2023-10-13 · Abstract (retrieved 2026-10-10)
- [6]
Google DeepMind demonstrated Gemini Robotics 2 on Apptronik's Apollo 2 humanoid, which walked to a table, picked up a watering can and placed it on a shelf from a spoken instruction. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 (retrieved 2026-10-10)
- [7]
RT-2 showed improved generalization to unfamiliar objects, could interpret commands not present in its robot training data, and performed rudimentary reasoning. confirmedas of 2023-07-28
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · arXiv (Google DeepMind authors) · 2023-07-28 · Abstract (retrieved 2026-10-10)
- [8]
OpenVLA (2024) is an open-source 7-billion-parameter vision-language-action model trained on 970,000 real-world robot demonstrations from Open X-Embodiment. confirmedas of 2024-06-13
- OpenVLA: An Open-Source Vision-Language-Action Model · arXiv · 2024-06-13 · Abstract (retrieved 2026-10-10)
- [9]
OpenVLA reported 16.5% higher absolute task success than the closed 55-billion-parameter RT-2-X across 29 tasks, and released its checkpoints and code. confirmedas of 2024-06-13
- OpenVLA: An Open-Source Vision-Language-Action Model · arXiv · 2024-06-13 · Abstract (retrieved 2026-10-10)
- [10]
Figure's Helix (February 2025) is a vision-language-action model with a slower onboard vision-language "System 2" at 7-9 Hz and a fast visuomotor "System 1" that outputs actions at 200 Hz. confirmedas of 2025-02-20
- Helix: A Vision-Language-Action Model for Generalist Humanoid Control · Figure AI · 2025-02-20 (retrieved 2026-10-10)
- [11]
GR00T N1 (March 2025) is an open humanoid foundation model with a vision-language module for understanding and a diffusion-transformer module that generates motor actions in real time. confirmedas of 2025-03-18
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots · arXiv (NVIDIA authors) · 2025-03-18 · Abstract (retrieved 2026-10-10)
- [12]
In March 2025 Google DeepMind introduced Gemini Robotics, a vision-language-action model built on Gemini 2.0, together with Gemini Robotics-ER, an embodied-reasoning model with enhanced spatial understanding. confirmedas of 2025-03-25
- Gemini Robotics: Bringing AI into the Physical World · arXiv (Google DeepMind authors) · 2025-03-25 · Abstract (retrieved 2026-10-10)
- [13]
Google DeepMind says Gemini Robotics 1.5 (September 2025) thinks before acting and learns across embodiments, while Gemini Robotics-ER 1.5 plans multi-step tasks and calls digital tools; ER 1.5 was made available to developers through the Gemini API. confirmedas of 2025-09-25
- Gemini Robotics 1.5 brings AI agents into the physical world · Google DeepMind · 2025-09-25 (retrieved 2026-10-10)
- Gemini Robotics 1.5 brings AI agents into the physical world · Google DeepMind · 2025-09-25 (retrieved 2026-10-10)
- [14]
In June 2025 Google DeepMind released Gemini Robotics On-Device, a VLA for bi-arm robots that runs locally without a data network and can be adapted to new tasks with 50 to 100 demonstrations. confirmedas of 2025-06-24
- Gemini Robotics On-Device brings AI to local robotic devices · Google DeepMind · 2025-06-24 (retrieved 2026-10-10)
- Gemini Robotics On-Device brings AI to local robotic devices · Google DeepMind · 2025-06-24 (retrieved 2026-10-10)
- [15]
Physical Intelligence's π0.5 (April 2025) uses co-training on heterogeneous tasks and multimodal examples (images, language commands, object detections, semantic subtask predictions and low-level actions) to perform long-horizon tasks such as cleaning a kitchen or bedroom in homes not seen during training. confirmedas of 2025-04-22
- π0.5: a Vision-Language-Action Model with Open-World Generalization · arXiv (Physical Intelligence authors) · 2025-04-22 · Abstract (retrieved 2026-10-10)
- [16]
In November 2025 Physical Intelligence introduced π*0.6, trained with a reinforcement-learning method called Recap that combines demonstrations, corrections and autonomous experience; the company says it more than doubled throughput on some of the hardest tasks. confirmedas of 2025-11-17
- π*0.6: a VLA that Learns from Experience · Physical Intelligence · 2025-11-17 (retrieved 2026-10-10)
- [17]
Physical Intelligence said in April 2026 that its π0.7 model performs dexterous tasks as well as fine-tuned specialists and shows first signs of compositional generalization, such as folding laundry on a robot with no laundry-folding data. confirmedas of 2026-04-16
- π0.7: a Steerable Model with Emergent Capabilities · Physical Intelligence · 2026-04-16 (retrieved 2026-10-10)
- [18]
On 5 January 2026 NVIDIA released GR00T N1.6, an open reasoning VLA for humanoids that unlocks full-body control and uses NVIDIA's Cosmos Reason model. confirmedas of 2026-01-05
- NVIDIA Releases New Physical AI Models as Global Partners Unveil Next-Generation Robots · NVIDIA · 2026-01-05 (retrieved 2026-10-10)
- [19]
As of October 2026 GR00T N1.7 was generally available under the Apache 2.0 licence, with a new vision-language backbone and 20,000 hours of human video in pretraining; NVIDIA says it performs comparably to N1.6 with better generalization and language following. confirmedas of 2026-10-10
- NVIDIA/Isaac-GR00T: NVIDIA Isaac GR00T N1.7 (README) · NVIDIA (GitHub) (retrieved 2026-10-10)
- [20]
On March 16, 2026 NVIDIA made GR00T N1.7 available in early access with commercial licensing and previewed GR00T N2, which it says succeeds at new tasks in new environments more than twice as often as leading VLA models and is slated for release by the end of 2026. confirmedas of 2026-03-16
- NVIDIA and Global Robotics Leaders Take Physical AI to the Real World · NVIDIA Newsroom · 2026-03-16 (retrieved 2026-10-10)
- [21]
Figure says Helix 02 (January 2026) controls the whole robot from pixels, unloading and reloading a dishwasher across a kitchen in a four-minute autonomous task, using a learned whole-body controller, System 0, trained on more than 1,000 hours of human motion data and sim-to-real reinforcement learning. confirmedas of 2026-01-27
- Introducing Helix 02: Full-Body Autonomy · Figure AI · 2026-01-27 (retrieved 2026-10-10)
- [22]
On September 17, 2026 Figure reported that Helix 2.5 performed tidying, towel folding and bed making in 30 Bay Area homes where no data had been collected, with 56% zero-shot success versus 9% without pretraining on its Index human-behavior dataset. confirmedas of 2026-09-17
- Helix 2.5: Zero-Shot 30-Home Generalization · Figure AI · 2026-09-17 (retrieved 2026-10-10)
- [23]
Hugging Face's SmolVLA (June 2025) is a small VLA trained on community-collected data that can be trained on a single GPU and deployed on consumer GPUs or CPUs, and its authors report performance comparable to VLAs ten times larger. confirmedas of 2025-06-02
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics · arXiv (Hugging Face authors) · 2025-06-02 (retrieved 2026-10-10)
- [24]
Physical Intelligence's FAST tokenizer (January 2025) compresses robot actions with the discrete cosine transform; combined with π0 it matched diffusion VLAs on 10,000 hours of robot data while cutting training time by up to five times. confirmedas of 2025-01-16
- FAST: Efficient Action Tokenization for Vision-Language-Action Models · arXiv (Physical Intelligence authors) · 2025-01-16 (retrieved 2026-10-10)
- [25]
NVIDIA describes GR00T N1.7 as a vision-language foundation model combined with a diffusion-transformer head that denoises continuous actions, adaptable to specific robots through post-training. confirmedas of 2026-10-10
- NVIDIA/Isaac-GR00T: NVIDIA Isaac GR00T N1.7 (README) · NVIDIA (GitHub) (retrieved 2026-10-10)
- [26]
Google DeepMind reported Gemini Robotics 2 success rates of 45.7% to 76.3% on general whole-body manipulation evaluations and 32% to 92% on multi-finger dexterity tasks. confirmedas of 2026-07-30
- Gemini Robotics 2 brings whole body intelligence to robots · Google DeepMind · 2026-07-30 · Results section (retrieved 2026-10-10)
- [27]
IEEE Spectrum reported that industrial customers expect about 99.99% reliability and that ISO safety standards for dynamically balancing legged robots were still being developed in 2025. reportedas of 2025-09-11
- Reality Is Ruining the Humanoid Robot Hype · IEEE Spectrum · 2025-09-11 (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Vision-language-action models: how robot foundation models work." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/vision-language-action-models
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerRobotics and embodied AI in 2026: a crash courseA crash course on robotics and embodied AI as of October 2026: how robots learn, robot foundation models, humanoids, self-driving and the open debates.
- ExplainerHow robots learn: demonstrations, trial and error, and dataA plain guide to how modern robots learn skills from human demonstrations, reinforcement learning and pooled datasets, and why data is the bottleneck.
- WikiGemini RoboticsGemini Robotics is Google DeepMind's family of robot AI models. What it does, how it evolved from RT-2 to Gemini Robotics 2, and its limits.
- WikiPhysical IntelligencePhysical Intelligence (π) is a startup building general-purpose robot foundation models such as π0 and π0.5. Its models, data approach and funding.
- WikiNVIDIA Isaac GR00TIsaac GR00T is NVIDIA's platform of open humanoid robot foundation models, simulation tools and reference hardware. Versions, data and partners.
- WikiFigure AIFigure AI builds general-purpose humanoid robots and its own Helix AI model. Its robots, BMW deployments, funding and 2026 milestones.