Explainer
How frontier AI capabilities are measured
Frontier AI is measured mostly with benchmarks, fixed sets of tasks scored automatically, plus newer measures such as how long a task an agent can finish on its own.[1] Benchmarks keep saturating: models scored over 90% on the once-standard MMLU test, which prompted harder tests such as Humanity's Last Exam.[2]
To follow frontier AI you need to read benchmark claims critically. Labs announce new models with tables of scores, and independent groups run their own tests.[3][4] This page explains the main kinds of measurement in use in 2026 and their limits.
Benchmarks and saturation
A benchmark is a fixed exam for AI: the same questions for every model, marked automatically. The problem is that models keep acing them. By 2025, models were scoring over 90% on MMLU, a widely used knowledge test, so researchers built Humanity’s Last Exam: 2,500 hard questions across dozens of subjects.[2] By September 2026 Anthropic reported 67.7% on it for Claude Opus 5.5 with tools.[3]
Static question-answer benchmarks saturate.[2] Humanity’s Last Exam was introduced in January 2025 with 2,500 closed-ended, expert-validated questions because frontier models exceeded 90% on MMLU.[2] Google reported 37.5% without tools for Gemini 3 Pro in November 2025,[5] and Anthropic reported 67.7% with tools for Claude Opus 5.5 in September 2026.[3] Scores with and without tools are reported separately, so the two figures are not like for like.[3]
Tests of real work
Newer benchmarks use real tasks. swe-bench gives an AI a real issue from an open-source project on GitHub and asks it to change the code to resolve it.[6] In 2023 the best model fixed under 2% of them;[7] by 2026 the AI Index reported near 100% on a cleaned-up version called SWE-bench Verified.[8] OSWorld does the same for using a desktop computer.[9]
SWE-bench draws 2,294 problems from GitHub issues and pull requests in 12 Python repositories;[6] SWE-bench Verified is a 500-task human-filtered subset.[10] As Verified nears saturation, labs also report harder variants such as SWE-bench Pro, where Anthropic gave Claude Mythos Preview 77.8%.[11] Agentic suites such as OSWorld, with 369 desktop tasks, test multi-step computer use.[9]
Easy for humans, hard for AI
The arc-agi benchmarks are built from tasks that people can solve. The newest version, ARC-AGI-3, is a set of games with no instructions: the player must work out the rules and the goal.[12] At launch in March 2026, people scored 100% and AI 0.51%.[13]
ARC-AGI-3 is interactive: agents must explore, infer goals, build world models and learn continuously,[14] and 100% means beating every game as efficiently as humans.[15] In September 2026 the ARC Prize Foundation reported 62.7% for GPT-6 Astra with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness.[16]
Time horizons
The evaluation group METR asks a different question: how long a task, in human working hours, can an AI finish on its own about half the time?[1] That number doubled about every seven months from 2019 to early 2025.[17]
The 50% time horizon is the human-expert task length at which a model succeeds half the time.[1] METR’s Time Horizon 1.1 estimated about 320 minutes for Claude Opus 4.5 and a doubling time of about 89 days since 2024,[18][4] but warned that confidence intervals were very wide and that few long tasks had human baselines.[19] Its estimate of at least 16 hours for an early Claude Mythos Preview sits at the top of what the suite can measure.[20]
Reading results carefully
Three cautions apply. Capabilities are uneven, so a high score in one area says little about another.[21] Benchmarks can diverge from real productivity, as METR’s developer trials showed.[22] And disclosure is shrinking: the 2026 AI Index reports that average Foundation Model Transparency Index scores fell to 40 from 58.[23]
Questions readers ask
What is benchmark saturation?
When top models score so high on a test that it no longer separates them. Humanity's Last Exam was created because models were scoring over 90% on popular benchmarks such as MMLU.[2]
What does a "time horizon" of 16 hours mean?
METR's 50% time horizon is the length of task, measured by how long human experts take, that a model completes with 50% success. A 16-hour horizon means tasks that take experts about 16 hours.[1][20]
Is there a test that humans find easy but AI finds hard?
ARC-AGI-3 was designed that way. At its March 2026 launch humans scored 100% and frontier AI 0.51%, though OpenAI's GPT-6 Astra reached 62.7% in September 2026.[13][16]
Do benchmark scores show real-world usefulness?
Not reliably. METR found experienced developers took 19% longer on real tasks when allowed to use early-2025 AI tools.[22]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [2]
Humanity's Last Exam, released in January 2025, is a benchmark of 2,500 expert-written questions across dozens of subjects, created because models were scoring over 90% on popular benchmarks such as MMLU. confirmedas of 2025-01-24
- Humanity's Last Exam · arXiv (Phan et al.) · 2025-01-24 · Abstract (retrieved 2026-10-10)
- Humanity's Last Exam · Center for AI Safety and Scale AI (retrieved 2026-10-10)
- [3]
Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 · Benchmark table (retrieved 2026-10-10)
- [4]
In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [5]
Google launched Gemini 3 on 18 November 2025, reporting 37.5% on Humanity's Last Exam without tools and 91.9% on GPQA Diamond for Gemini 3 Pro. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Benchmark section (retrieved 2026-10-10)
- [6]
SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (the arXiv page renders the counts in LaTeX as $2,294$ and $12$) (retrieved 2026-10-10)
- SWE-bench Lite · SWE-bench (retrieved 2026-10-10)
- [7]
When SWE-bench was introduced in 2023, the best-performing model, Claude 2, solved 1.96% of the issues. confirmedas of 2023-10-10
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (retrieved 2026-10-10)
- [8]
The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [9]
OSWorld, a 2024 benchmark of 369 real computer tasks across operating systems, found humans succeeded on 72.36% of tasks against 12.24% for the best AI model at the time. confirmedas of 2024-04-11
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · arXiv (Xie et al.) · 2024-04-11 · Abstract (retrieved 2026-10-10)
- [10]
As of October 2026 the SWE-bench project runs several leaderboards, including SWE-bench Verified (500 human-filtered tasks), Lite, Multilingual and Multimodal. confirmedas of 2026-10-10
- SWE-bench Verified · SWE-bench (retrieved 2026-10-10)
- SWE-bench leaderboards · SWE-bench · Site navigation (Verified, Lite, Multilingual, Multimodal leaderboards) (retrieved 2026-10-10)
- [11]
Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 · Benchmark table (retrieved 2026-10-10)
- [12]
ARC-AGI-3 consists of game-like environments with no instructions, rules or stated goals, which an agent must explore to work out how they function. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · What ARC-AGI-3 measures (retrieved 2026-10-10)
- [13]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [14]
The ARC Prize Foundation describes ARC-AGI-3 as an interactive reasoning benchmark in which AI agents must explore novel environments, acquire goals on the fly, build adaptable world models and learn continuously. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [15]
On ARC-AGI-3, a 100% score means an AI agent can beat every game as efficiently as humans. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [16]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [17]
METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [18]
Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [19]
METR cautioned that its Time Horizon 1.1 confidence intervals were still very wide and that only 5 of its 31 long tasks had human baseline measurements. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [20]
METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10
- METR says it can barely measure Claude Mythos · The Decoder · 2026-05-10 (retrieved 2026-10-10)
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates log, 8 May 2026 entry: Added Claude Mythos Preview (early); the point estimate appears in the page's chart data (retrieved 2026-10-10)
- [21]
The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [22]
In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · arXiv (METR) · Abstract (retrieved 2026-10-10)
- [23]
The 2026 AI Index reports that average scores on the Foundation Model Transparency Index fell to 40 from 58 the year before, with the most capable models disclosing the least. confirmedas of 2026-04-01
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
Revision history (1)
- Page created.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"How frontier AI capabilities are measured." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-ai-capabilities-are-measured
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- ExplainerHow AI agents workWhat an AI agent is and how it works: language models using tools in a loop, computer use, MCP connectors and long tasks.
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.
- WikiMETRMETR is an AI evaluation group known for measuring how long a task AI agents can complete, and for trials of AI's effect on developers.