Skip to content
ContentLora

    Tip: press / anywhere to search.

    Analysis

    How fast are AI agents really improving?

    On benchmarks, AI agents are improving very fast: METR estimated in January 2026 that the length of software tasks agents can complete was doubling about every 89 days since 2024.[1] Evidence that this translates into real-world productivity is much weaker,[2] and the best models are now outrunning the tests built to measure them.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligence
    Show:

    Whether AI agents are close to doing long stretches of skilled work on their own is one of the main open questions in frontier AI. The benchmark evidence points one way, the productivity evidence is murkier, and measurement itself is under strain.[1][2][3]

    The benchmark evidence

    METR’s 2025 study found the 50% time horizon of frontier models doubled about every seven months from 2019, with Claude 3.7 Sonnet at around 50 minutes.[4] Its January 2026 revision estimated a doubling time of about 131 days since 2023 and about 89 days since 2024, with Claude Opus 4.5 at about 320 minutes.[1][5] METR then estimated at least 16 hours for an early version of Claude Mythos Preview tested in March 2026.[3] The 2026 AI Index reports agents went from about 12% to about 66% on OSWorld and that SWE-bench Verified scores rose from 60% to near 100% in one year.[6][7] On ARC-AGI-3, frontier AI scored 0.51% at launch in March 2026 and GPT-6 Astra 62.7% in September.[8][9]

    The productivity evidence

    In METR’s 2025 randomized trial, experienced open-source developers took 19% longer on tasks when allowed early-2025 AI tools, despite expecting a 24% speedup.[10] A larger late-2025 follow-up estimated speedups of -18% and -4% for two groups, with confidence intervals spanning zero, and was skewed because many developers declined to work without AI.[11] METR concluded that true gains were likely much higher than measured but that its data was very weak evidence.[2] The International AI Safety Report 2026 notes that systems can excel at hard tasks while failing at simpler ones.[12]

    The two bodies of evidence measure different things. Time horizons use self-contained tasks with clear success criteria; real work involves unclear goals, messy codebases and human review. A model can double its benchmark horizon without a matching change in office productivity, and the reverse is also possible: developers refusing to work without AI is itself a signal of perceived value, even if it ruins the experiment.

    Measurement under strain

    METR says measurements above 16 hours are unreliable with its current task suite,[3] and in Time Horizon 1.1 only 5 of 31 long tasks had human baselines.[13] Its public time-horizons page was marked as no longer actively updated on 8 September 2026.[14] On ARC-AGI-3, the same GPT-6 Astra model scored 62.7% with one harness and 99.9% with another.[9] The 2026 AI Index reports falling transparency scores among foundation model developers.[15]

    Fewer independent measurements, more lab-reported scores and large harness effects together make it harder for outsiders to check claims about the frontier. Developer-reported figures, such as Anthropic’s 93.9% on SWE-bench Verified for Mythos Preview, are useful but cannot replace independent evaluation.[16]

    What may happen next

    METR’s 2025 paper projected that, if trends held, AI could automate many month-long software tasks within five years.[17] With measured doubling times now nearer three months than seven, that horizon would arrive sooner if the trend continues, but the measurements behind the faster rate have wide uncertainty. A reasonable expectation for 2027 is that evaluators publish new, longer task suites and that the debate shifts from “can agents do it” to “how reliably, at what cost, and with how much oversight”. Low confidence.

    The International AI Safety Report 2026 sums up the uncertainty: progress through 2030 is uncertain, but current trends are consistent with continued improvement.[18]

    Competing views

    Fast and accelerating

    Time horizons, coding and computer-use scores have all jumped, and the doubling time has shortened since 2024. On this view agents capable of multi-day work are arriving on schedule.[1][3][6][9]

    Benchmarks overstate real gains

    Benchmarks are narrow and saturate. Controlled trials have not shown clear productivity gains, capabilities remain uneven, and labs disclose less about their models.[10][2][12][15]

    Progress is real, measurement is failing

    Capabilities are rising, but evaluators can no longer measure the frontier reliably, and scores depend heavily on the harness around the model.[13][14][9]

    Questions readers ask

    How quickly are AI agents' capabilities growing?

    METR found the length of software tasks agents can complete doubled about every seven months from 2019 to early 2025, and in January 2026 estimated about 89 days since 2024.[4][1]

    Do AI coding tools make developers faster?

    The evidence is unclear. METR's 2025 trial found a 19% slowdown with early-2025 tools; its follow-up said true gains were likely much higher but its evidence was very weak.[10][2]

    Why is it getting harder to measure agent progress?

    The best models exceed what test suites can measure. METR says measurements above 16 hours are unreliable with its current tasks, and SWE-bench Verified scores are near 100%.[3][7]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29

    2. [2]

      METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24

    3. [3]

      METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10

    4. [4]

      METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18

    5. [5]

      Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29

    6. [6]

      The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01

    7. [7]

      The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01

    8. [8]

      The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25

      • Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
    9. [9]

      On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03

    10. [10]

      In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12

    11. [11]

      METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24

    12. [12]

      The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24

    13. [13]

      METR cautioned that its Time Horizon 1.1 confidence intervals were still very wide and that only 5 of its 31 long tasks had human baseline measurements. confirmedas of 2026-01-29

    14. [14]

      METR's public time-horizons page carried a notice, in its 8 September 2026 update, that it is no longer actively updated. confirmedas of 2026-09-08

    15. [15]

      The 2026 AI Index reports that average scores on the Foundation Model Transparency Index fell to 40 from 58 the year before, with the most capable models disclosing the least. confirmedas of 2026-04-01

    16. [16]

      Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07

    17. [17]

      METR's 2025 study projected that if the trend continued, within five years AI systems would be able to automate many software tasks that take humans a month. reportedas of 2025-03-18· forecast

    18. [18]

      The International AI Safety Report 2026 concludes that the trajectory of AI progress through 2030 is uncertain but that current trends are consistent with continued improvement. confirmedas of 2026-02-24

    Revision history (1)
    1. Page created.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "How fast are AI agents really improving?." ContentLora, updated Oct 10, 2026. https://contentlora.com/analysis/ai-agent-progress-debate

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.