Analysis
How fast are AI agents really improving?
On benchmarks, AI agents are improving very fast: METR estimated in January 2026 that the length of software tasks agents can complete was doubling about every 89 days since 2024.[1] Evidence that this translates into real-world productivity is much weaker,[2] and the best models are now outrunning the tests built to measure them.[3]
Whether AI agents are close to doing long stretches of skilled work on their own is one of the main open questions in frontier AI. The benchmark evidence points one way, the productivity evidence is murkier, and measurement itself is under strain.[1][2][3]
The benchmark evidence
METR’s 2025 study found the 50% time horizon of frontier models doubled about every seven months from 2019, with Claude 3.7 Sonnet at around 50 minutes.[4] Its January 2026 revision estimated a doubling time of about 131 days since 2023 and about 89 days since 2024, with Claude Opus 4.5 at about 320 minutes.[1][5] METR then estimated at least 16 hours for an early version of Claude Mythos Preview tested in March 2026.[3] The 2026 AI Index reports agents went from about 12% to about 66% on OSWorld and that SWE-bench Verified scores rose from 60% to near 100% in one year.[6][7] On ARC-AGI-3, frontier AI scored 0.51% at launch in March 2026 and GPT-6 Astra 62.7% in September.[8][9]
The productivity evidence
In METR’s 2025 randomized trial, experienced open-source developers took 19% longer on tasks when allowed early-2025 AI tools, despite expecting a 24% speedup.[10] A larger late-2025 follow-up estimated speedups of -18% and -4% for two groups, with confidence intervals spanning zero, and was skewed because many developers declined to work without AI.[11] METR concluded that true gains were likely much higher than measured but that its data was very weak evidence.[2] The International AI Safety Report 2026 notes that systems can excel at hard tasks while failing at simpler ones.[12]
The two bodies of evidence measure different things. Time horizons use self-contained tasks with clear success criteria; real work involves unclear goals, messy codebases and human review. A model can double its benchmark horizon without a matching change in office productivity, and the reverse is also possible: developers refusing to work without AI is itself a signal of perceived value, even if it ruins the experiment.
Measurement under strain
METR says measurements above 16 hours are unreliable with its current task suite,[3] and in Time Horizon 1.1 only 5 of 31 long tasks had human baselines.[13] Its public time-horizons page was marked as no longer actively updated on 8 September 2026.[14] On ARC-AGI-3, the same GPT-6 Astra model scored 62.7% with one harness and 99.9% with another.[9] The 2026 AI Index reports falling transparency scores among foundation model developers.[15]
Fewer independent measurements, more lab-reported scores and large harness effects together make it harder for outsiders to check claims about the frontier. Developer-reported figures, such as Anthropic’s 93.9% on SWE-bench Verified for Mythos Preview, are useful but cannot replace independent evaluation.[16]
What may happen next
METR’s 2025 paper projected that, if trends held, AI could automate many month-long software tasks within five years.[17] With measured doubling times now nearer three months than seven, that horizon would arrive sooner if the trend continues, but the measurements behind the faster rate have wide uncertainty. A reasonable expectation for 2027 is that evaluators publish new, longer task suites and that the debate shifts from “can agents do it” to “how reliably, at what cost, and with how much oversight”. Low confidence.
The International AI Safety Report 2026 sums up the uncertainty: progress through 2030 is uncertain, but current trends are consistent with continued improvement.[18]
Competing views
Fast and accelerating
Time horizons, coding and computer-use scores have all jumped, and the doubling time has shortened since 2024. On this view agents capable of multi-day work are arriving on schedule.[1][3][6][9]
Questions readers ask
How quickly are AI agents' capabilities growing?
METR found the length of software tasks agents can complete doubled about every seven months from 2019 to early 2025, and in January 2026 estimated about 89 days since 2024.[4][1]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [2]
METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [3]
METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10
- METR says it can barely measure Claude Mythos · The Decoder · 2026-05-10 (retrieved 2026-10-10)
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates log, 8 May 2026 entry: Added Claude Mythos Preview (early); the point estimate appears in the page's chart data (retrieved 2026-10-10)
- [4]
METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [5]
Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [6]
The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [7]
The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [8]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [9]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [10]
In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · arXiv (METR) · Abstract (retrieved 2026-10-10)
- [11]
METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [12]
The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [13]
METR cautioned that its Time Horizon 1.1 confidence intervals were still very wide and that only 5 of its 31 long tasks had human baseline measurements. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [14]
METR's public time-horizons page carried a notice, in its 8 September 2026 update, that it is no longer actively updated. confirmedas of 2026-09-08
- Task-Completion Time Horizons of Frontier AI Models · METR (retrieved 2026-10-10)
- [15]
The 2026 AI Index reports that average scores on the Foundation Model Transparency Index fell to 40 from 58 the year before, with the most capable models disclosing the least. confirmedas of 2026-04-01
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
- [16]
Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 · Benchmark table (retrieved 2026-10-10)
- [17]
METR's 2025 study projected that if the trend continued, within five years AI systems would be able to automate many software tasks that take humans a month. reportedas of 2025-03-18· forecast
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [18]
The International AI Safety Report 2026 concludes that the trajectory of AI progress through 2030 is uncertain but that current trends are consistent with continued improvement. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
Revision history (1)
- Page created.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"How fast are AI agents really improving?." ContentLora, updated Oct 10, 2026. https://contentlora.com/analysis/ai-agent-progress-debate
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow AI agents workWhat an AI agent is and how it works: language models using tools in a loop, computer use, MCP connectors and long tasks.
- ExplainerHow frontier AI capabilities are measuredHow researchers measure what frontier AI can do: benchmarks like SWE-bench, Humanity's Last Exam and ARC-AGI, and METR time horizons.
- DevelopingFrontier AI tracker: reasoning models and agentsA dated timeline of frontier AI milestones in reasoning models and AI agents, from o1 and DeepSeek-R1 to GPT-6, Claude Mythos and Gemini 4.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.