Skip to content
ContentLora

    Tip: press / anywhere to search.

    concept

    ARC-AGI

    Also known as ARC-AGI-3, ARC-AGI-2, ARC Prize

    ARC-AGI is a series of benchmarks run by the ARC Prize Foundation. Its newest version, ARC-AGI-3, is an interactive set of games in which AI agents must explore, discover goals and learn as they go.[1] At launch in March 2026 humans scored 100% and frontier AI 0.51%;[2] in September 2026 OpenAI's GPT-6 Astra scored 62.7%.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligenceScience
    Key facts

    ARC-AGI benchmarks aim to measure how efficiently an AI system learns something new, rather than what it already knows.[1] The ARC Prize Foundation states that as long as there is a gap between AI and human learning, AGI has not been achieved.[4] That focus differs from knowledge-heavy tests such as Humanity’s Last Exam, which uses expert-written academic questions.[5]

    Origins: ARC-AGI-1

    The first benchmark in the series, now called ARC-AGI-1, was introduced in 2019 by François Chollet in his paper “On the Measure of Intelligence”.[6] It consists of 800 grid-based puzzles; each gives only a few example input-output pairs, so the solver must infer the rule on the spot.[7] From 2019 until late 2024 no AI system solved it.[8] According to the ARC Prize Foundation, scores went from 0% with GPT-3 in 2020 to just 5% with GPT-4o in 2024.[9]

    That changed with reasoning models. In December 2024 the Foundation reported that a preview of OpenAI‘s o3 scored 75.7% on ARC-AGI-1’s semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute.[10] The gap between those two numbers is why the Foundation reports cost alongside score; see test-time-compute.[10]

    ARC-AGI-3: games without instructions

    ARC-AGI-3, launched on 25 March 2026, is interactive and turn-based.[2] It is a set of game-like environments with no instructions, rules or stated goals; an agent must explore each one, work out how it functions and discover what winning looks like.[11] The Foundation describes it as testing whether agents can acquire goals on the fly, build adaptable world models and learn continuously.[1] Scores account for efficiency: 100% means beating every game as efficiently as humans.[12]

    From 0.51% to 62.7%

    At launch, humans scored 100% and frontier AI 0.51%.[2] On 3 September 2026 the ARC Prize Foundation reported that OpenAI’s GPT-6 Astra scored 62.7% on the semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness.[3] With the adapted harness, Astra used fewer actions than the human baseline on 96.0% of levels.[13] The two results used the same model with different harnesses, the software that connects a model to the games.[3] The Foundation called Astra’s result a noticeable step-function change in frontier model capabilities.[3]

    ARC-AGI-2

    ARC-AGI-2, launched in 2025, was built to stress-test reasoning systems, and the Foundation says log-linear scaling of compute is not enough to beat it.[14] Every task in its public evaluation set was solved by at least two people within two attempts.[15] In November 2025 Google reported that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2.[16] The Foundation has said it will keep running its Grand Prize until a high-efficiency, open-source solution scores 85%.[17] The ARC-AGI-2 Grand Prize remains part of ARC Prize 2026, reserved for the best open-source solution.[18]

    ARC Prize 2026

    ARC Prize 2026 offers more than $2 million in total across an ARC-AGI-3 agent competition hosted on Kaggle and the ARC-AGI-2 Grand Prize.[18] The Foundation also reports what runs cost, such as $26,098 for Astra’s standard-harness result.[3]

    How it fits with other measures

    Other widely used tests measure different abilities. swe-bench uses real software issues from GitHub,[19] and METR‘s time horizons measure how long a task, in human-expert time, an agent completes half the time.[20] ARC-AGI-3 instead gives agents no instructions or stated goals, so they must discover what to do.[11]

    Questions readers ask

    What makes ARC-AGI-3 different from other benchmarks?

    It is interactive. Agents play game-like environments with no instructions, rules or stated goals, and must work out how each one works.[11]

    What does a 100% score on ARC-AGI-3 mean?

    That an AI agent can beat every game as efficiently as humans.[12]

    How well does AI do on ARC-AGI-3?

    Frontier AI scored 0.51% at launch in March 2026. In September 2026 the ARC Prize Foundation reported 62.7% for GPT-6 Astra with its standard harness and 99.9% with a provider-adapted harness.[2][3]

    Is there prize money?

    Yes. ARC Prize 2026 offers more than $2 million across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution.[18]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      The ARC Prize Foundation describes ARC-AGI-3 as an interactive reasoning benchmark in which AI agents must explore novel environments, acquire goals on the fly, build adaptable world models and learn continuously. confirmedas of 2026-10-10

      • ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
    2. [2]

      The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25

      • Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
    3. [3]

      On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03

    4. [4]

      The ARC Prize Foundation states that as long as there is a gap between AI and human learning, AGI has not been achieved. confirmedas of 2026-10-10

      • ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
    5. [5]

      Humanity's Last Exam, released in January 2025, is a benchmark of 2,500 expert-written questions across dozens of subjects, created because models were scoring over 90% on popular benchmarks such as MMLU. confirmedas of 2025-01-24

    6. [6]

      The first ARC benchmark, now called ARC-AGI-1, was introduced in 2019 by François Chollet in his paper "On the Measure of Intelligence". confirmedas of 2026-10-10

      • ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
    7. [7]

      ARC-AGI-1 consists of 800 grid-based reasoning puzzles, each typically giving only around three example input-output pairs. confirmedas of 2026-10-10

      • ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
    8. [8]

      From its introduction in 2019 until late 2024, ARC-AGI remained unsolved by AI systems. confirmedas of 2026-10-10

      • ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
    9. [9]

      According to the ARC Prize Foundation, ARC-AGI-1 scores went from 0% with GPT-3 in 2020 to 5% with GPT-4o in 2024. confirmedas of 2024-12-20

    10. [10]

      In December 2024 the ARC Prize Foundation reported that a preview of OpenAI's o3 scored 75.7% on the ARC-AGI-1 semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute. confirmedas of 2024-12-20

    11. [11]

      ARC-AGI-3 consists of game-like environments with no instructions, rules or stated goals, which an agent must explore to work out how they function. confirmedas of 2026-03-25

      • Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · What ARC-AGI-3 measures (retrieved 2026-10-10)
    12. [12]

      On ARC-AGI-3, a 100% score means an AI agent can beat every game as efficiently as humans. confirmedas of 2026-10-10

      • ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
    13. [13]

      The ARC Prize Foundation reported that, with the provider-adapted harness, GPT-6 Astra used fewer actions than the human baseline on 96.0% of ARC-AGI-3 levels. confirmedas of 2026-09-03

    14. [14]

      ARC-AGI-2, launched in 2025, is designed to stress-test state-of-the-art AI reasoning systems, and the ARC Prize Foundation says log-linear scaling of compute is insufficient to beat it. confirmedas of 2026-10-10

      • ARC-AGI-2 · ARC Prize Foundation (retrieved 2026-10-10)
    15. [15]

      Every task in ARC-AGI-2's public evaluation set was solved by at least two humans within two attempts. confirmedas of 2026-10-10

      • ARC-AGI-2 · ARC Prize Foundation (retrieved 2026-10-10)
    16. [16]

      Google reported in November 2025 that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2 and 41.0% on Humanity's Last Exam. confirmedas of 2025-11-18

    17. [17]

      The ARC Prize Foundation has said it will run its Grand Prize until a high-efficiency, open-source solution scores 85%. confirmedas of 2024-12-20

    18. [18]

      ARC Prize 2026 offers more than $2 million in prizes across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution. confirmedas of 2026-03-25

    19. [19]

      SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11

    20. [20]

      METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18

    Revision history (2)
    1. Page created.
    2. Added the history of ARC-AGI-1 and ARC-AGI-2, OpenAI o3's December 2024 result and the 85% Grand Prize target.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "ARC-AGI." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/arc-agi

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.