concept
ARC-AGI
Also known as ARC-AGI-3, ARC-AGI-2, ARC Prize
ARC-AGI is a series of benchmarks run by the ARC Prize Foundation. Its newest version, ARC-AGI-3, is an interactive set of games in which AI agents must explore, discover goals and learn as they go.[1] At launch in March 2026 humans scored 100% and frontier AI 0.51%;[2] in September 2026 OpenAI's GPT-6 Astra scored 62.7%.[3]
Key facts
ARC-AGI benchmarks aim to measure how efficiently an AI system learns something new, rather than what it already knows.[1] The ARC Prize Foundation states that as long as there is a gap between AI and human learning, AGI has not been achieved.[4] That focus differs from knowledge-heavy tests such as Humanity’s Last Exam, which uses expert-written academic questions.[5]
Origins: ARC-AGI-1
The first benchmark in the series, now called ARC-AGI-1, was introduced in 2019 by François Chollet in his paper “On the Measure of Intelligence”.[6] It consists of 800 grid-based puzzles; each gives only a few example input-output pairs, so the solver must infer the rule on the spot.[7] From 2019 until late 2024 no AI system solved it.[8] According to the ARC Prize Foundation, scores went from 0% with GPT-3 in 2020 to just 5% with GPT-4o in 2024.[9]
That changed with reasoning models. In December 2024 the Foundation reported that a preview of OpenAI‘s o3 scored 75.7% on ARC-AGI-1’s semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute.[10] The gap between those two numbers is why the Foundation reports cost alongside score; see test-time-compute.[10]
ARC-AGI-3: games without instructions
ARC-AGI-3, launched on 25 March 2026, is interactive and turn-based.[2] It is a set of game-like environments with no instructions, rules or stated goals; an agent must explore each one, work out how it functions and discover what winning looks like.[11] The Foundation describes it as testing whether agents can acquire goals on the fly, build adaptable world models and learn continuously.[1] Scores account for efficiency: 100% means beating every game as efficiently as humans.[12]
From 0.51% to 62.7%
At launch, humans scored 100% and frontier AI 0.51%.[2] On 3 September 2026 the ARC Prize Foundation reported that OpenAI’s GPT-6 Astra scored 62.7% on the semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness.[3] With the adapted harness, Astra used fewer actions than the human baseline on 96.0% of levels.[13] The two results used the same model with different harnesses, the software that connects a model to the games.[3] The Foundation called Astra’s result a noticeable step-function change in frontier model capabilities.[3]
ARC-AGI-2
ARC-AGI-2, launched in 2025, was built to stress-test reasoning systems, and the Foundation says log-linear scaling of compute is not enough to beat it.[14] Every task in its public evaluation set was solved by at least two people within two attempts.[15] In November 2025 Google reported that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2.[16] The Foundation has said it will keep running its Grand Prize until a high-efficiency, open-source solution scores 85%.[17] The ARC-AGI-2 Grand Prize remains part of ARC Prize 2026, reserved for the best open-source solution.[18]
ARC Prize 2026
ARC Prize 2026 offers more than $2 million in total across an ARC-AGI-3 agent competition hosted on Kaggle and the ARC-AGI-2 Grand Prize.[18] The Foundation also reports what runs cost, such as $26,098 for Astra’s standard-harness result.[3]
How it fits with other measures
Other widely used tests measure different abilities. swe-bench uses real software issues from GitHub,[19] and METR‘s time horizons measure how long a task, in human-expert time, an agent completes half the time.[20] ARC-AGI-3 instead gives agents no instructions or stated goals, so they must discover what to do.[11]
Questions readers ask
What makes ARC-AGI-3 different from other benchmarks?
It is interactive. Agents play game-like environments with no instructions, rules or stated goals, and must work out how each one works.[11]
What does a 100% score on ARC-AGI-3 mean?
That an AI agent can beat every game as efficiently as humans.[12]
How well does AI do on ARC-AGI-3?
Frontier AI scored 0.51% at launch in March 2026. In September 2026 the ARC Prize Foundation reported 62.7% for GPT-6 Astra with its standard harness and 99.9% with a provider-adapted harness.[2][3]
Is there prize money?
Yes. ARC Prize 2026 offers more than $2 million across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution.[18]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
The ARC Prize Foundation describes ARC-AGI-3 as an interactive reasoning benchmark in which AI agents must explore novel environments, acquire goals on the fly, build adaptable world models and learn continuously. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [2]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [3]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [4]
The ARC Prize Foundation states that as long as there is a gap between AI and human learning, AGI has not been achieved. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [5]
Humanity's Last Exam, released in January 2025, is a benchmark of 2,500 expert-written questions across dozens of subjects, created because models were scoring over 90% on popular benchmarks such as MMLU. confirmedas of 2025-01-24
- Humanity's Last Exam · arXiv (Phan et al.) · 2025-01-24 · Abstract (retrieved 2026-10-10)
- Humanity's Last Exam · Center for AI Safety and Scale AI (retrieved 2026-10-10)
- [6]
The first ARC benchmark, now called ARC-AGI-1, was introduced in 2019 by François Chollet in his paper "On the Measure of Intelligence". confirmedas of 2026-10-10
- ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
- [7]
ARC-AGI-1 consists of 800 grid-based reasoning puzzles, each typically giving only around three example input-output pairs. confirmedas of 2026-10-10
- ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
- [8]
From its introduction in 2019 until late 2024, ARC-AGI remained unsolved by AI systems. confirmedas of 2026-10-10
- ARC-AGI-1 · ARC Prize Foundation (retrieved 2026-10-10)
- [9]
According to the ARC Prize Foundation, ARC-AGI-1 scores went from 0% with GPT-3 in 2020 to 5% with GPT-4o in 2024. confirmedas of 2024-12-20
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub · ARC Prize Foundation · 2024-12-20 (retrieved 2026-10-10)
- [10]
In December 2024 the ARC Prize Foundation reported that a preview of OpenAI's o3 scored 75.7% on the ARC-AGI-1 semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute. confirmedas of 2024-12-20
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub · ARC Prize Foundation · 2024-12-20 (retrieved 2026-10-10)
- [11]
ARC-AGI-3 consists of game-like environments with no instructions, rules or stated goals, which an agent must explore to work out how they function. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · What ARC-AGI-3 measures (retrieved 2026-10-10)
- [12]
On ARC-AGI-3, a 100% score means an AI agent can beat every game as efficiently as humans. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [13]
The ARC Prize Foundation reported that, with the provider-adapted harness, GPT-6 Astra used fewer actions than the human baseline on 96.0% of ARC-AGI-3 levels. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [14]
ARC-AGI-2, launched in 2025, is designed to stress-test state-of-the-art AI reasoning systems, and the ARC Prize Foundation says log-linear scaling of compute is insufficient to beat it. confirmedas of 2026-10-10
- ARC-AGI-2 · ARC Prize Foundation (retrieved 2026-10-10)
- [15]
Every task in ARC-AGI-2's public evaluation set was solved by at least two humans within two attempts. confirmedas of 2026-10-10
- ARC-AGI-2 · ARC Prize Foundation (retrieved 2026-10-10)
- [16]
Google reported in November 2025 that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2 and 41.0% on Humanity's Last Exam. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Deep Think section (retrieved 2026-10-10)
- [17]
The ARC Prize Foundation has said it will run its Grand Prize until a high-efficiency, open-source solution scores 85%. confirmedas of 2024-12-20
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub · ARC Prize Foundation · 2024-12-20 (retrieved 2026-10-10)
- [18]
ARC Prize 2026 offers more than $2 million in prizes across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Prize details (retrieved 2026-10-10)
- [19]
SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (the arXiv page renders the counts in LaTeX as $2,294$ and $12$) (retrieved 2026-10-10)
- SWE-bench Lite · SWE-bench (retrieved 2026-10-10)
- [20]
METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Added the history of ARC-AGI-1 and ARC-AGI-2, OpenAI o3's December 2024 result and the 85% Grand Prize target.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"ARC-AGI." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/arc-agi
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow frontier AI capabilities are measuredHow researchers measure what frontier AI can do: benchmarks like SWE-bench, Humanity's Last Exam and ARC-AGI, and METR time horizons.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.
- DevelopingFrontier AI tracker: reasoning models and agentsA dated timeline of frontier AI milestones in reasoning models and AI agents, from o1 and DeepSeek-R1 to GPT-6, Claude Mythos and Gemini 4.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.
- WikiMETRMETR is an AI evaluation group known for measuring how long a task AI agents can complete, and for trials of AI's effect on developers.
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.