Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    Vision-language-action models: how robot foundation models work

    A vision-language-action (VLA) model is a neural network that takes camera images and a language instruction and outputs robot actions. It is usually built on a pretrained vision-language model.[1][2] Since RT-2 in 2023, VLAs have become the main approach to general-purpose robot control. As of 2026, models from Google DeepMind, NVIDIA, Physical Intelligence and Figure control whole humanoid bodies.[3][4]

    Editor reviewedUpdated Robotics and embodied AIArtificial intelligence

    Vision-language-action models are the robotics version of the “foundation model” idea. One large network, pretrained on broad data, is adapted to control robots across many tasks.[1][5] They are the core technology behind most 2025-2026 humanoid demonstrations.[3][4]

    The basic idea

    A chatbot reads words and writes words. A VLA looks through the robot’s cameras, reads an instruction such as “put the watering can in the green bin”, and writes out motor commands instead of text.[6] It starts from a model that already understands images and language from web data. That helps it handle objects and commands it never saw during robot training.[7]

    Google DeepMind’s RT-2 (2023) co-fine-tuned a vision-language model on web and robot data by expressing robot actions as text tokens. The resulting model generalized to novel objects and showed rudimentary semantic reasoning.[1][7] OpenVLA (2024) is an open 7-billion-parameter version, trained on 970,000 demonstrations from Open X-Embodiment. It reported 16.5% higher absolute success than the 55-billion-parameter RT-2-X across 29 tasks.[8][9]

    Fast hands, slow thinking

    Robots must move smoothly many times a second, but big language models are slow. Many VLAs therefore split the work in two. A slower “thinking” part understands the scene and the instruction, and a fast “moving” part turns that into motion.[10][11]

    Discrete action tokens limit control frequency and precision, so newer models attach continuous action heads. π0 uses flow matching on top of a pretrained VLM.[2] NVIDIA’s GR00T N1 pairs a vision-language module with a diffusion transformer that generates actions in real time.[11] Figure’s Helix runs an onboard VLM (“System 2”) at 7-9 Hz that conditions a visuomotor policy (“System 1”) outputting actions at 200 Hz.[10] Google DeepMind separates an embodied-reasoning model (Gemini Robotics-ER) from the action model.[12]

    Main models as of October 2026

    • Gemini Robotics (gemini-robotics): Google DeepMind‘s family. Gemini Robotics 1.5 (September 2025) “thinks” before acting, an On-Device version runs locally on bi-arm robots, and Gemini Robotics 2 (July 2026) is the first in the family to control whole humanoids, including walking.[13][14][3]
    • π series (physical-intelligence): π0 (2024) and π0.5 (2025), which cleaned kitchens and bedrooms in homes not seen during training.[2][15] π*0.6 (November 2025) added reinforcement learning from the robot’s own experience, and π0.7 (April 2026) showed early signs of combining skills in new ways.[16][17]
    • GR00T (isaac-gr00t): NVIDIA‘s open humanoid models. N1.6 (January 2026) added full-body control; N1.7 is generally available under Apache 2.0; N2 is due by the end of 2026.[18][19][20]
    • Helix (figure-ai): Figure’s in-house VLA. Helix 02 (January 2026) controls the whole body, and Helix 2.5 was tested zero-shot in 30 homes.[21][22]
    • Open research models: OpenVLA (7 billion parameters) and Hugging Face’s SmolVLA, which runs on consumer GPUs or even CPUs.[8][23]

    Turning motions into tokens

    Language models predict tokens, so a VLA needs a way to express continuous motions as tokens or to generate them directly. RT-2 wrote actions as text tokens.[1] Physical Intelligence’s FAST tokenizer compresses action sequences with the discrete cosine transform and, with π0, matched diffusion-based models while training up to five times faster.[24] Others skip tokens: GR00T uses a diffusion-transformer head that denoises continuous actions.[25]

    Limits

    Reported success rates show how far these systems still are from industrial reliability. Gemini Robotics 2 scored 45.7% to 76.3% on general whole-body manipulation evaluations,[26] and Figure’s Helix 2.5 reached 56% zero-shot success in unseen homes.[22] Industrial customers are reported to expect around 99.99% reliability.[27] These figures come from the developers’ own evaluations, so they are not directly comparable across companies.[26][22]

    Questions readers ask

    What does "vision-language-action" mean?

    The model takes images (vision) and an instruction (language) as input and outputs robot motor commands (action). RT-2 did this by encoding the actions as text tokens.[1]

    Are there open-source VLAs?

    Yes. OpenVLA is an open 7-billion-parameter model trained on 970,000 robot demonstrations. NVIDIA's GR00T N1 was released as an open humanoid foundation model.[8][11]

    How well do the newest VLAs work?

    They are far from perfect. Google DeepMind reported Gemini Robotics 2 success rates of 45.7% to 76.3% on whole-body manipulation tests, and Figure reported 56% zero-shot success in unseen homes.[26][22]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      Google DeepMind's RT-2 (July 2023) trained a single vision-language model to output robot actions by expressing those actions as text tokens. confirmedas of 2023-07-28

    2. [2]

      Physical Intelligence's π0 (October 2024) is a vision-language-action model that adds a flow-matching action generator on top of a pretrained vision-language model. confirmedas of 2024-10-31

    3. [3]

      On July 30, 2026 Google DeepMind announced Gemini Robotics 2, a vision-language-action model it says can control full humanoids, including walking and manipulation, as well as bi-arm robots; previous models in the family controlled a humanoid's upper body for tabletop tasks. confirmedas of 2026-07-30

    4. [4]

      Figure says Helix 02 coordinates Figure 03's hands, arms, torso and feet. confirmedas of 2026-06-30

    5. [5]

      The RT-X models trained on Open X-Embodiment (paper first posted October 2023) showed positive transfer, improving the capabilities of multiple robots by using experience from other robot platforms. confirmedas of 2023-10-13

    6. [6]

      Google DeepMind demonstrated Gemini Robotics 2 on Apptronik's Apollo 2 humanoid, which walked to a table, picked up a watering can and placed it on a shelf from a spoken instruction. confirmedas of 2026-07-30

    7. [7]

      RT-2 showed improved generalization to unfamiliar objects, could interpret commands not present in its robot training data, and performed rudimentary reasoning. confirmedas of 2023-07-28

    8. [8]

      OpenVLA (2024) is an open-source 7-billion-parameter vision-language-action model trained on 970,000 real-world robot demonstrations from Open X-Embodiment. confirmedas of 2024-06-13

    9. [9]

      OpenVLA reported 16.5% higher absolute task success than the closed 55-billion-parameter RT-2-X across 29 tasks, and released its checkpoints and code. confirmedas of 2024-06-13

    10. [10]

      Figure's Helix (February 2025) is a vision-language-action model with a slower onboard vision-language "System 2" at 7-9 Hz and a fast visuomotor "System 1" that outputs actions at 200 Hz. confirmedas of 2025-02-20

    11. [11]

      GR00T N1 (March 2025) is an open humanoid foundation model with a vision-language module for understanding and a diffusion-transformer module that generates motor actions in real time. confirmedas of 2025-03-18

    12. [12]

      In March 2025 Google DeepMind introduced Gemini Robotics, a vision-language-action model built on Gemini 2.0, together with Gemini Robotics-ER, an embodied-reasoning model with enhanced spatial understanding. confirmedas of 2025-03-25

    13. [13]

      Google DeepMind says Gemini Robotics 1.5 (September 2025) thinks before acting and learns across embodiments, while Gemini Robotics-ER 1.5 plans multi-step tasks and calls digital tools; ER 1.5 was made available to developers through the Gemini API. confirmedas of 2025-09-25

    14. [14]

      In June 2025 Google DeepMind released Gemini Robotics On-Device, a VLA for bi-arm robots that runs locally without a data network and can be adapted to new tasks with 50 to 100 demonstrations. confirmedas of 2025-06-24

    15. [15]

      Physical Intelligence's π0.5 (April 2025) uses co-training on heterogeneous tasks and multimodal examples (images, language commands, object detections, semantic subtask predictions and low-level actions) to perform long-horizon tasks such as cleaning a kitchen or bedroom in homes not seen during training. confirmedas of 2025-04-22

    16. [16]

      In November 2025 Physical Intelligence introduced π*0.6, trained with a reinforcement-learning method called Recap that combines demonstrations, corrections and autonomous experience; the company says it more than doubled throughput on some of the hardest tasks. confirmedas of 2025-11-17

    17. [17]

      Physical Intelligence said in April 2026 that its π0.7 model performs dexterous tasks as well as fine-tuned specialists and shows first signs of compositional generalization, such as folding laundry on a robot with no laundry-folding data. confirmedas of 2026-04-16

    18. [18]

      On 5 January 2026 NVIDIA released GR00T N1.6, an open reasoning VLA for humanoids that unlocks full-body control and uses NVIDIA's Cosmos Reason model. confirmedas of 2026-01-05

    19. [19]

      As of October 2026 GR00T N1.7 was generally available under the Apache 2.0 licence, with a new vision-language backbone and 20,000 hours of human video in pretraining; NVIDIA says it performs comparably to N1.6 with better generalization and language following. confirmedas of 2026-10-10

    20. [20]

      On March 16, 2026 NVIDIA made GR00T N1.7 available in early access with commercial licensing and previewed GR00T N2, which it says succeeds at new tasks in new environments more than twice as often as leading VLA models and is slated for release by the end of 2026. confirmedas of 2026-03-16

    21. [21]

      Figure says Helix 02 (January 2026) controls the whole robot from pixels, unloading and reloading a dishwasher across a kitchen in a four-minute autonomous task, using a learned whole-body controller, System 0, trained on more than 1,000 hours of human motion data and sim-to-real reinforcement learning. confirmedas of 2026-01-27

    22. [22]

      On September 17, 2026 Figure reported that Helix 2.5 performed tidying, towel folding and bed making in 30 Bay Area homes where no data had been collected, with 56% zero-shot success versus 9% without pretraining on its Index human-behavior dataset. confirmedas of 2026-09-17

    23. [23]

      Hugging Face's SmolVLA (June 2025) is a small VLA trained on community-collected data that can be trained on a single GPU and deployed on consumer GPUs or CPUs, and its authors report performance comparable to VLAs ten times larger. confirmedas of 2025-06-02

    24. [24]

      Physical Intelligence's FAST tokenizer (January 2025) compresses robot actions with the discrete cosine transform; combined with π0 it matched diffusion VLAs on 10,000 hours of robot data while cutting training time by up to five times. confirmedas of 2025-01-16

    25. [25]

      NVIDIA describes GR00T N1.7 as a vision-language foundation model combined with a diffusion-transformer head that denoises continuous actions, adaptable to specific robots through post-training. confirmedas of 2026-10-10

    26. [26]

      Google DeepMind reported Gemini Robotics 2 success rates of 45.7% to 76.3% on general whole-body manipulation evaluations and 32% to 92% on multi-finger dexterity tasks. confirmedas of 2026-07-30

    27. [27]

      IEEE Spectrum reported that industrial customers expect about 99.99% reliability and that ISO safety standards for dynamically balancing legged robots were still being developed in 2025. reportedas of 2025-09-11

    Revision history (2)
    1. Page created.
    2. Updated the model list (Gemini Robotics 1.5 and On-Device, GR00T N1.6 and N1.7, π*0.6 and π0.7, Helix 02) and added open VLAs and action tokenization.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Vision-language-action models: how robot foundation models work." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/vision-language-action-models

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.