Skip to content
ContentLora

    Tip: press / anywhere to search.

    technology

    Open X-Embodiment

    Also known as OXE, RT-X, Open X-Embodiment Dataset

    Open X-Embodiment is an open robot-learning dataset and set of models released in 2023. It pooled more than a million real robot trajectories from 22 robot types, collected by 21 institutions.[1] Its RT-X models showed that training on data from many robots can improve each robot,[2] an idea that shaped later open models such as OpenVLA.[3]

    Editor reviewedUpdated Robotics and embodied AIArtificial intelligence
    Key facts

    Open X-Embodiment combined robot data from dozens of labs into one shared dataset that anyone can train on.[1] It tested whether robots, like language models, get better when trained on broad, mixed data.[2]

    Background

    The project grew out of Google’s Robotics Transformer work. The RT-1 paper (December 2022) argued that general robot policies need open-ended, task-agnostic training on large and diverse real-robot data, with models big enough to absorb it.[4] Open X-Embodiment applied that idea across institutions, with Google DeepMind researchers training the RT-X models on the pooled data.[5]

    What it contains

    The collaboration pooled 60 existing robot datasets from 34 research labs.[6] The result is more than a million real robot trajectories from 22 robot embodiments, covering 527 skills and 160,266 tasks.[1] “Embodiment” means the robot’s physical form, such as a single arm, two arms or a mobile base.

    The RT-X result

    The team trained two model families. RT-1-X outperformed models trained on individual datasets by 50% in the small-data setting. RT-2-X, based on Google DeepMind’s RT-2 vision-language-action model, did about three times better on emergent-skill evaluations.[5] The paper’s central finding was positive transfer: one model trained on many robots’ data improved several robots at once.[2]

    Influence

    OpenVLA (2024) is an open 7-billion-parameter model trained on 970,000 Open X-Embodiment demonstrations. It reported 16.5% higher absolute success than the 55-billion-parameter RT-2-X across 29 tasks, and it released its weights and code.[3][7] Industry models have also adopted the multi-robot, multi-source approach. Physical Intelligence‘s π0 trains on several robot types, and NVIDIA’s GR00T mixes robot data with human video and synthetic data.[8][9]

    Other open projects built on the same base. Octo (2024), an open generalist policy, was trained on 800,000 Open X-Embodiment trajectories and could be fine-tuned to new sensors and action spaces in a few hours on consumer GPUs.[10] DROID (2024) added 76,000 demonstrations, or 350 hours, collected in 564 scenes by 50 people on three continents, and released everything openly.[11] Such shared platforms let companies’ open models be tested outside their labs: Physical Intelligence’s openpi release included π0 versions fine-tuned for the DROID and ALOHA setups.[12] Hugging Face’s SmolVLA (2025) went small: trained on community-collected data, it runs on consumer GPUs or even CPUs, and its authors report performance comparable to models ten times larger.[13]

    Limits

    About a million trajectories is still small next to the text used to train language models. Goldberg describes this shortfall as a “100,000-year data gap”.[14] Companies such as Figure now build their own large datasets for pretraining.[15] Figure said in August 2026 that its Index app had collected over 16 million videos and that it planned to spend over $1 billion on data and compute within a year.[16] Tools are also being shared across datasets: Physical Intelligence released FAST+, an action tokenizer trained on one million real robot trajectories, for use with many robot types.[17]

    Questions readers ask

    What is Open X-Embodiment?

    It is a pooled dataset of more than a million real robot trajectories from 22 robot types, assembled by 21 institutions, plus the RT-X models trained on it.[1]

    What did it show?

    A high-capacity model trained on the mixed data showed positive transfer, improving the capabilities of several robots by learning from others' experience.[2]

    Which models use it?

    OpenVLA, an open 7-billion-parameter VLA, was trained on 970,000 demonstrations from Open X-Embodiment.[3]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      The Open X-Embodiment dataset pooled more than one million real robot trajectories from 22 robot embodiments, collected through a collaboration of 21 institutions and covering 527 skills. confirmedas of 2025-05-14

    2. [2]

      The RT-X models trained on Open X-Embodiment (paper first posted October 2023) showed positive transfer, improving the capabilities of multiple robots by using experience from other robot platforms. confirmedas of 2023-10-13

    3. [3]

      OpenVLA (2024) is an open-source 7-billion-parameter vision-language-action model trained on 970,000 real-world robot demonstrations from Open X-Embodiment. confirmedas of 2024-06-13

    4. [4]

      Google's RT-1 Robotics Transformer (December 2022) argued that open-ended, task-agnostic training on large and diverse real-robot data with high-capacity models is key to general robot policies. confirmedas of 2022-12-13

    5. [5]

      RT-1-X outperformed models trained on individual datasets by 50% in the small-data domain, and RT-2-X showed about three times better performance on emergent-skill evaluations. confirmedas of 2026-10-10

    6. [6]

      Open X-Embodiment was assembled from 60 existing robot datasets from 34 robotics research labs. confirmedas of 2026-10-10

    7. [7]

      OpenVLA reported 16.5% higher absolute task success than the closed 55-billion-parameter RT-2-X across 29 tasks, and released its checkpoints and code. confirmedas of 2024-06-13

    8. [8]

      π0 was trained on data from multiple robot types, including single-arm, dual-arm and mobile manipulators, and evaluated on tasks such as laundry folding, table cleaning and box assembly. confirmedas of 2024-10-31

    9. [9]

      GR00T N1 was trained on a mixture of real-robot trajectories, human videos and synthetically generated data, and was deployed on the Fourier GR-1 humanoid. confirmedas of 2025-03-18

    10. [10]

      Octo (2024), an open-source generalist robot policy, was trained on 800,000 trajectories from Open X-Embodiment and could be fine-tuned to new sensors and action spaces in a few hours on consumer GPUs. confirmedas of 2024-05-20

    11. [11]

      The DROID dataset (2024) contains 76,000 demonstration trajectories, or 350 hours of interaction, collected across 564 scenes and 84 tasks by 50 collectors in North America, Asia and Europe over 12 months, and is fully open source. confirmedas of 2024-03-19

    12. [12]

      Physical Intelligence's openpi release included π0 checkpoints fine-tuned for widely available research platforms such as ALOHA and DROID. confirmedas of 2025-02-04

    13. [13]

      Hugging Face's SmolVLA (June 2025) is a small VLA trained on community-collected data that can be trained on a single GPU and deployed on consumer GPUs or CPUs, and its authors report performance comparable to VLAs ten times larger. confirmedas of 2025-06-02

    14. [14]

      UC Berkeley roboticist Ken Goldberg argued in August 2025 Science Robotics papers that robots face a "100,000-year data gap" compared with the text used to train language models. confirmedas of 2025-08-27

    15. [15]

      On September 17, 2026 Figure reported that Helix 2.5 performed tidying, towel folding and bed making in 30 Bay Area homes where no data had been collected, with 56% zero-shot success versus 9% without pretraining on its Index human-behavior dataset. confirmedas of 2026-09-17

    16. [16]

      In August 2026 Figure launched Index, an app that pays people to record videos for robot training; Figure said it had paid $15 million to contributors, received over 16 million videos and committed to spend over $1 billion on data and compute in the next 12 months. confirmedas of 2026-08-25

    17. [17]

      Physical Intelligence released FAST+, a universal robot action tokenizer trained on one million real robot action trajectories. confirmedas of 2025-01-16

    Revision history (2)
    1. Page created.
    2. Added the RT-1 lineage and the open models and datasets that followed (Octo, DROID, SmolVLA, FAST+).

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Open X-Embodiment." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/open-x-embodiment

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.