organization
METR
Also known as Model Evaluation and Threat Research
METR is a research group that evaluates frontier AI systems. It is best known for the "time horizon" measure, the length of task in human-expert time that a model completes half the time,[1] which it found doubled about every seven months from 2019 to early 2025.[2] In 2026 it estimated at least 16 hours for an early Claude Mythos Preview.[3]
Key facts
What METR is
METR (Model Evaluation and Threat Research) describes itself as a research nonprofit that measures whether and when AI systems might threaten catastrophic harm to society.[4][5] It evaluates frontier models to inform the public about their risks and capabilities, mostly by testing how far they can carry out substantial tasks on their own.[6] It says it is funded by donations and has not accepted funding from AI companies, although it uses significant free tokens from them.[7]
Its evaluations cover frontier models from major labs: for example, METR found that OpenAI‘s GPT-5.1-Codex-Max did not pose significant catastrophic risks via AI self-improvement or rogue replication.[8] It also prototyped the responsible scaling policy approach, which it says nine AI developers have adopted; see frontier-safety-frameworks.[9] In August 2026 the UK AI Security Institute said it intended to work with METR on an independent review of an incident in which AI agents took unsanctioned real-world actions during a government cyber evaluation.[5]
Influence
METR’s measurements feed into major assessments of AI progress: the International AI Safety Report 2026 cites the seven-month doubling estimate.[10] The group also runs randomized trials of how AI tools affect real developers.[11]
The time-horizon measure
METR’s 2025 paper defined the 50%-task-completion time horizon: the human-expert duration of tasks a model completes with 50% success.[1] It put Claude 3.7 Sonnet at around 50 minutes and found the frontier horizon had doubled about every seven months since 2019.[2] The paper attributed the growth mainly to better reliability, recovery from mistakes, logical reasoning and tool use,[12] and projected that, if the trend held, within five years AI could automate many software tasks that take humans a month.[13]
Time Horizon 1.1 and Mythos
In January 2026 METR released Time Horizon 1.1, expanding its suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31.[14] It estimated a doubling time of about 131 days since 2023 and about 89 days since 2024, and horizons of about 320 minutes for Claude Opus 4.5, 214 for GPT-5 and 121 for o3.[14][15] METR warned that confidence intervals were very wide and that only 5 of 31 long tasks had human baselines.[16] In May 2026 it reported at least 16 hours, with a 95% confidence interval of 8.5 to 55 hours, for an early version of Claude Mythos Preview tested in March, at the upper end of what it can measure.[3] Its own page logged the addition with a warning that measurements above 16 hours are unreliable with its current task suite.[3] Its public time-horizons page noted on 8 September 2026 that it is no longer actively updated.[17]
Developer productivity trials
In a 2025 randomized controlled trial, 16 experienced open-source developers completed 246 tasks; allowing early-2025 AI tools increased completion time by 19%, although the developers had expected a 24% speedup.[11] A larger follow-up with 57 developers and more than 800 tasks estimated speedups of -18% for returning developers and -4% for new ones, but METR said a growing number of developers declined to take part because they did not want to work without AI.[18] It concluded that true gains were likely much higher than measured, that the data was very weak evidence, and that it would change its study design.[19]
Why it matters
METR’s work links two questions the field argues about: how capable agents are on benchmark-style tasks, and how much they help in practice.[14][19] Both are covered in the analysis of agent progress.[10]
Questions readers ask
What is a METR time horizon?
The length of task, measured by how long human experts take, that an AI model completes with 50% success.[1]
How fast are time horizons growing?
METR's 2025 study found doubling about every seven months since 2019. Its January 2026 revision estimated about 131 days since 2023 and about 89 days since 2024.[2][14]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [2]
METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [3]
METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10
- METR says it can barely measure Claude Mythos · The Decoder · 2026-05-10 (retrieved 2026-10-10)
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates log, 8 May 2026 entry: Added Claude Mythos Preview (early); the point estimate appears in the page's chart data (retrieved 2026-10-10)
- [4]
METR describes itself as a research nonprofit that measures whether and when AI systems might threaten catastrophic harm to society. confirmedas of 2026-10-10
- About METR · METR (retrieved 2026-10-10)
- [5]
AISI said it intended to work with METR on an independent third-party review of the August 2026 incident. confirmedas of 2026-08-04
- Incident Report: unsanctioned agent behaviour during cyber testing · UK AI Security Institute · 2026-08-04 (retrieved 2026-10-10)
- [6]
METR evaluates frontier AI models to inform the public about their risks and capabilities, mostly by assessing how far AI systems can autonomously carry out substantial tasks. confirmedas of 2026-10-10
- About METR · METR (retrieved 2026-10-10)
- [7]
METR says it is funded by donations and has not accepted funding from AI companies, although it uses significant free tokens from them. confirmedas of 2026-10-10
- About METR · METR (retrieved 2026-10-10)
- [8]
METR evaluated OpenAI's GPT-5.1-Codex-Max and found it did not pose significant catastrophic risks via AI self-improvement or rogue replication. confirmedas of 2026-10-10
- About METR · METR (retrieved 2026-10-10)
- [9]
METR says it prototyped the Responsible Scaling Policies approach, which it says has been adopted by nine AI developers. confirmedas of 2026-10-10
- About METR · METR (retrieved 2026-10-10)
- [10]
The International AI Safety Report 2026 cites estimates that the complexity of software tasks AI agents can accomplish doubles roughly every seven months. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [11]
In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · arXiv (METR) · Abstract (retrieved 2026-10-10)
- [12]
METR attributed the growth in time horizons mainly to better reliability, ability to recover from mistakes, logical reasoning and tool use. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [13]
METR's 2025 study projected that if the trend continued, within five years AI systems would be able to automate many software tasks that take humans a month. reportedas of 2025-03-18· forecast
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [14]
In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [15]
Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [16]
METR cautioned that its Time Horizon 1.1 confidence intervals were still very wide and that only 5 of its 31 long tasks had human baseline measurements. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [17]
METR's public time-horizons page carried a notice, in its 8 September 2026 update, that it is no longer actively updated. confirmedas of 2026-09-08
- Task-Completion Time Horizons of Frontier AI Models · METR (retrieved 2026-10-10)
- [18]
METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [19]
METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"METR." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/metr
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.
- ExplainerHow frontier AI capabilities are measuredHow researchers measure what frontier AI can do: benchmarks like SWE-bench, Humanity's Last Exam and ARC-AGI, and METR time horizons.
- DevelopingFrontier AI tracker: reasoning models and agentsA dated timeline of frontier AI milestones in reasoning models and AI agents, from o1 and DeepSeek-R1 to GPT-6, Claude Mythos and Gemini 4.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.