Skip to content
ContentLora

    Tip: press / anywhere to search.

    product

    Gemini Robotics

    Also known as Gemini Robotics 2, Gemini Robotics-ER, Gemini Robotics On-Device

    Gemini Robotics is Google DeepMind's family of vision-language-action and embodied-reasoning models for controlling robots. It was first released in March 2025, built on Gemini 2.0.[1] Gemini Robotics 2, announced on July 30, 2026, is the first in the family to control whole humanoids, including walking.[2]

    Editor reviewedUpdated Robotics and embodied AIArtificial intelligence
    Key facts

    Gemini Robotics is Google DeepMind‘s main effort to bring its Gemini models into the physical world. It is one of the best-documented examples of a vision-language-action (VLA) model.[1][2]

    Origins

    Google DeepMind’s RT-2 (2023) showed that a vision-language model could output robot actions encoded as text tokens.[3] RT-2 generalized better to unfamiliar objects, could interpret commands that were not in its robot training data and showed rudimentary reasoning, the pattern later Gemini models built on.[4] Its RT-2 architecture was later trained on the multi-lab Open X-Embodiment dataset as RT-2-X.[5] In March 2025 it introduced Gemini Robotics, a VLA built on Gemini 2.0, and Gemini Robotics-ER, a model for embodied reasoning such as spatial understanding.[1] The paper reported that the model could learn new tasks from about 100 demonstrations and adapt to new robot designs.[6] In June 2025 Google DeepMind added Gemini Robotics On-Device, a VLA for two-armed robots that runs on the robot itself without a data network. Developers could adapt it to new tasks with 50 to 100 demonstrations.[7] A companion Gemini Robotics SDK let trusted testers evaluate the model on their own tasks and try it in the MuJoCo physics simulator before using real hardware.[8]

    Gemini Robotics 1.5 followed in September 2025.[9] Google DeepMind says the 1.5 VLA “thinks” before acting and shows its reasoning, and that it learns across different robot bodies. Its companion, Gemini Robotics-ER 1.5, plans multi-step tasks and can call digital tools. ER 1.5 was opened to developers through the Gemini API, while the VLA stayed with select partners.[10]

    Gemini Robotics 2

    On July 30, 2026, Google DeepMind released three models: the Gemini Robotics 2 VLA, the Gemini Robotics ER 2 reasoning model and Gemini Robotics On-Device 2.[11] Earlier versions controlled a humanoid’s upper body for tabletop tasks. Google DeepMind says the new VLA controls whole humanoids, including walking.[2] In one demonstration, Apptronik’s Apollo 2 humanoid walked to a table, picked up a watering can and placed it on a shelf.[12]

    Google DeepMind also introduced a safety benchmark, ASIMOV-Agentic. It says ER 2 is better at detecting nearby people and triggering safety tool calls.[13]

    Performance and limits

    Google DeepMind’s own evaluations show wide variation. Success ranged from 45.7% to 76.3% on general whole-body manipulation and from 32% to 92% on multi-finger dexterity tasks.[14] Access to the action models is limited to early-access partners, so independent tests are scarce.[11]

    Partners

    Google DeepMind lists Agile Robots, Apptronik and Boston Dynamics among its partners. It also runs a program with more than 100 trusted testers.[15] Waymo’s driving model also builds on Gemini.[16]

    How it compares

    Gemini Robotics is one of three big general-purpose robot model efforts, alongside NVIDIA’s open Isaac GR00T models and start-up Physical Intelligence‘s π series.[17][18] Unlike GR00T N1.7, which is released under the Apache 2.0 licence, Google DeepMind’s action models remain limited to partners and trusted testers.[17][11]

    Questions readers ask

    What is Gemini Robotics 2?

    It is Google DeepMind's July 2026 vision-language-action model. Google DeepMind says it can control full humanoids from walking to fine manipulation, as well as two-armed robots.[2]

    Can I use Gemini Robotics?

    The ER 2 reasoning model is available in Google AI Studio. The action (VLA) and on-device models are limited to early-access partners.[11]

    How reliable is it?

    Google DeepMind reported success rates of 45.7% to 76.3% on general whole-body manipulation evaluations and 32% to 92% on multi-finger dexterity tasks.[14]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      In March 2025 Google DeepMind introduced Gemini Robotics, a vision-language-action model built on Gemini 2.0, together with Gemini Robotics-ER, an embodied-reasoning model with enhanced spatial understanding. confirmedas of 2025-03-25

    2. [2]

      On July 30, 2026 Google DeepMind announced Gemini Robotics 2, a vision-language-action model it says can control full humanoids, including walking and manipulation, as well as bi-arm robots; previous models in the family controlled a humanoid's upper body for tabletop tasks. confirmedas of 2026-07-30

    3. [3]

      Google DeepMind's RT-2 (July 2023) trained a single vision-language model to output robot actions by expressing those actions as text tokens. confirmedas of 2023-07-28

    4. [4]

      RT-2 showed improved generalization to unfamiliar objects, could interpret commands not present in its robot training data, and performed rudimentary reasoning. confirmedas of 2023-07-28

    5. [5]

      RT-1-X outperformed models trained on individual datasets by 50% in the small-data domain, and RT-2-X showed about three times better performance on emergent-skill evaluations. confirmedas of 2026-10-10

    6. [6]

      The Gemini Robotics paper reports that the model can be adapted to new tasks with as few as about 100 demonstrations and to new robot designs. confirmedas of 2025-03-25

    7. [7]

      In June 2025 Google DeepMind released Gemini Robotics On-Device, a VLA for bi-arm robots that runs locally without a data network and can be adapted to new tasks with 50 to 100 demonstrations. confirmedas of 2025-06-24

    8. [8]

      Alongside Gemini Robotics On-Device, Google DeepMind released a Gemini Robotics SDK for trusted testers to evaluate the model on their own tasks and test it in the MuJoCo physics simulator. confirmedas of 2025-06-24

    9. [9]

      Google DeepMind released Gemini Robotics 1.5 in September 2025. confirmedas of 2026-07-30

    10. [10]

      Google DeepMind says Gemini Robotics 1.5 (September 2025) thinks before acting and learns across embodiments, while Gemini Robotics-ER 1.5 plans multi-step tasks and calls digital tools; ER 1.5 was made available to developers through the Gemini API. confirmedas of 2025-09-25

    11. [11]

      The Gemini Robotics 2 release has three models (the Gemini Robotics 2 VLA, the Gemini Robotics ER 2 reasoning model and Gemini Robotics On-Device 2). ER 2 is available in Google AI Studio, while the VLA and on-device models are limited to early-access partners. confirmedas of 2026-07-30

    12. [12]

      Google DeepMind demonstrated Gemini Robotics 2 on Apptronik's Apollo 2 humanoid, which walked to a table, picked up a watering can and placed it on a shelf from a spoken instruction. confirmedas of 2026-07-30

    13. [13]

      With Gemini Robotics 2, Google DeepMind introduced the ASIMOV-Agentic safety benchmark and said ER 2 can better detect nearby humans and trigger safety tool calls. confirmedas of 2026-07-30

    14. [14]

      Google DeepMind reported Gemini Robotics 2 success rates of 45.7% to 76.3% on general whole-body manipulation evaluations and 32% to 92% on multi-finger dexterity tasks. confirmedas of 2026-07-30

    15. [15]

      Google DeepMind lists Agile Robots, Apptronik and Boston Dynamics among its Gemini Robotics partners and runs a program with more than 100 trusted testers. confirmedas of 2026-10-10

    16. [16]

      Waymo describes a Waymo Foundation Model that combines a fast sensor-fusion encoder with a driving vision-language model built on Gemini, and that powers its driver, simulator and critic. confirmedas of 2025-12-09

    17. [17]

      As of October 2026 GR00T N1.7 was generally available under the Apache 2.0 licence, with a new vision-language backbone and 20,000 hours of human video in pretraining; NVIDIA says it performs comparably to N1.6 with better generalization and language following. confirmedas of 2026-10-10

    18. [18]

      Physical Intelligence said in April 2026 that its π0.7 model performs dexterous tasks as well as fine-tuned specialists and shows first signs of compositional generalization, such as folding laundry on a robot with no laundry-folding data. confirmedas of 2026-04-16

    19. [19]

      On March 16, 2026 NVIDIA made GR00T N1.7 available in early access with commercial licensing and previewed GR00T N2, which it says succeeds at new tasks in new environments more than twice as often as leading VLA models and is slated for release by the end of 2026. confirmedas of 2026-03-16

    20. [20]

      Physical Intelligence's π0.5 (April 2025) uses co-training on heterogeneous tasks and multimodal examples (images, language commands, object detections, semantic subtask predictions and low-level actions) to perform long-horizon tasks such as cleaning a kitchen or bedroom in homes not seen during training. confirmedas of 2025-04-22

    Revision history (2)
    1. Page created.
    2. Added Gemini Robotics On-Device (June 2025) and details of Gemini Robotics 1.5 and ER 1.5 (September 2025).

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Gemini Robotics." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/gemini-robotics

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.