● Developing story
Frontier AI tracker: reasoning models and agents
Since late 2024, frontier AI has shifted to reasoning models trained with reinforcement learning[1] and to agents that use tools and computers.[2] In 2026 the leading labs began gating their most capable models over cybersecurity risk,[3][4] while agent benchmarks such as ARC-AGI-3 moved from near zero to majority success.[5] In June 2026 the US briefly applied export controls to Anthropic's newest models.[6][7]
What we know
- Reasoning models are trained with large-scale reinforcement learning to think in a chain of thought[1]
- METR estimated agents' task horizons doubled about every 89 days since 2024[31]
- An early Claude Mythos Preview reached a time horizon of at least 16 hours[23]
- GPT-6 Astra scored 62.7% on ARC-AGI-3, up from 0.51% for frontier AI at launch[5]
- OpenAI rated GPT-6 Astra Critical for cybersecurity capability[4]
- MCP is governed by the Linux Foundation's Agentic AI Foundation[32]
- US and Chinese models have traded the top performance ranking since early 2025[33]
- OpenAI treats GPT-6.1 Sol as Critical for cybersecurity[11]
- Fable 5.1 is the same underlying model as Mythos 5.1, with added safeguards[15]
- CAISI estimated DeepSeek V4 Pro lagged the US frontier by about eight months[17]
What we don't know yet
- Whether Claude Mythos Preview or a successor will be made generally available, and under what safeguards
- How long a task the newest models can complete, now that METR's suite cannot measure above 16 hours reliably
- Whether benchmark gains translate into measurable productivity gains in real workplaces
- Whether any ARC Prize 2026 entry reaches human-level efficiency on ARC-AGI-3 under competition rules
- Whether reasoning-training scale-ups slow, as Epoch AI projected in 2025
Story status
This tracker follows the frontier of AI reasoning models and agents. It starts with the arrival of reinforcement-learning reasoning models in late 2024[1] and is updated as labs, evaluators and benchmark organisers publish new results.
Where things stand in October 2026
OpenAI, Anthropic and Google DeepMind each released new frontier models between September and early October 2026.[8][9][10] OpenAI followed Astra with GPT-6.1 Sol, which it also treats as Critical for cybersecurity,[11] and Anthropic completed its 5.5 family with Sonnet and Haiku.[12] Their most capable models are increasingly released in stages over cybersecurity risk: Anthropic has not made Claude Mythos Preview generally available,[3] OpenAI rated GPT-6 Astra Critical for cyber capability,[4] and Google first offered Gemini 4 Argon to cybersecurity professionals.[13] Anthropic still limits Mythos 5.1 to vetted organisations and sells the same model with extra safeguards as Fable 5.1.[14][15]
Open-weight competition continues through DeepSeek, which released V4.1-Flash in September and is phasing out V4-Pro.[16] The US government’s CAISI estimated in May that DeepSeek V4 Pro lagged the US frontier by about eight months.[17]
Gated models and government
Release decisions are increasingly shaped by cyber risk and by governments. In June 2026 a US export control directive forced Anthropic to switch off Fable 5 and Mythos 5 for all customers until the controls were lifted on 30 June.[6][7] Anthropic has since moved its gated cyber access into a tiered Cyber Verification Program.[18] For export rules on the chips that train these models, see US AI chip export controls and AI chip supply concentration.
Measuring a moving frontier
Government testers report fast gains on cyber tasks. The UK AI Security Institute estimated in February 2026 that the length of cyber tasks models could complete was doubling every 4.7 months, and said Claude Mythos Preview and GPT-5.5 then exceeded that trend.[19] It has also warned that evaluations with fixed compute budgets can systematically underestimate the newest agents, because the longest tasks are the first to be cut off.[20][21] That echoes an older lesson from ARC-AGI: in December 2024 an o3 preview scored 75.7% within a fixed compute limit and 87.5% with about 172 times more compute.[22] See test-time-compute and METR for how labs and evaluators measure these trends.
What to watch next
- Independent measurement. METR says it cannot reliably measure time horizons above 16 hours with its current tasks and has stopped actively updating its public page.[23][24] Watch for new, longer task suites.
- ARC Prize 2026. The competition offers more than $2 million; watch whether entries under competition rules approach the scores reported for GPT-6 Astra.[25][5]
- Gated models. Watch whether Mythos-class models reach general availability and how safeguards such as routing flagged queries to weaker models hold up.[26] Anthropic expects other companies to have Mythos-class models within 6 to 12 months of June 2026.[27]
- Real-world productivity. METR has said it will redesign its developer studies after inconclusive results.[28]
- Chain-of-thought oversight. Labs such as OpenAI now monitor full reasoning trajectories; whether chains of thought stay readable is an open question.[29][30]
Timeline
34 confirmed
confirmed
Anthropic launches its Cyber Mission[63]
confirmed
Anthropic releases Claude Haiku 5.5[12]
confirmed
OpenAI ships GPT-6 Sol and Luna in ChatGPT[67][68]
Both are rated High, below Critical, for cyber and bio capability; neither reaches High for AI self-improvement.
confirmed
Anthropic folds Glasswing into an expanded Cyber Verification Program[18][62]
Anthropic reported (company figures) that Glasswing partners found at least 129,000 verified vulnerabilities from April to July.
confirmed
Google DeepMind introduces Gemini 4 Argon, first for cyber defenders[10][13][66]
confirmed
OpenAI's GPT-6.1 Sol rated Critical for cyber[11]
confirmed
Anthropic releases Claude Sonnet 5.5[12]
confirmed
Anthropic releases Claude Opus 5.5[9][65]
Anthropic says it performs at Claude Fable 5.1 level on most work at 40% lower cost than Opus 5.
confirmed
DeepSeek releases V4.1-Flash and phases out V4-Pro[16]
confirmed
METR stops actively updating its time-horizons page[24]
confirmed
OpenAI releases GPT-6 Astra, rated Critical for cyber[8][4]
confirmed
GPT-6 Astra scores 62.7% on ARC-AGI-3[5][64]
confirmed
Anthropic releases Claude Fable 5.1 and Mythos 5.1[61][14][15]
confirmed
DeepSeek makes V4-Pro generally available[60]
confirmed
OpenAI publishes the GPT-5.6 system card[58][59]
OpenAI treated Sol, Terra and Luna as High, not Critical, for cyber and bio capability.
confirmed
US lifts export controls on Fable 5 and Mythos 5[7][57]
Fable 5 returned globally on 1 July; Mythos 5 was restored for a set of US organisations.
confirmed
Anthropic pulls Claude Fable 5 under a US export control directive[6]
Three days after release, Anthropic said a US government export control directive required it to disable Fable 5 and Mythos 5 for all customers.
confirmed
Anthropic releases Claude Fable 5 and Claude Mythos 5[56]
The two share one model; Mythos 5, with fewer safeguards, went only to Glasswing partners.
confirmed
Anthropic extends Project Glasswing to about 150 more organisations[55][27]
confirmed
METR measures Claude Mythos Preview at 16+ hours[54][23]
METR said the estimate was at the upper end of what its task suite can measure.
confirmed
DeepSeek releases V4 preview with open weights[52][53]
confirmed
Anthropic announces Claude Mythos Preview and Project Glasswing[50][51][3]
Anthropic withheld general release over cybersecurity risk and gave access to defenders of critical software.
confirmed
ARC-AGI-3 launches; frontier AI scores 0.51%[49][25]
confirmed
International AI Safety Report 2026 published[47][48]
The report highlights inference-time scaling and uneven, "jagged" capabilities.
confirmed
METR releases Time Horizon 1.1[31][46]
METR estimated a doubling time of about 89 days since 2024, with Claude Opus 4.5 at about 320 minutes.
confirmed
Linux Foundation forms the Agentic AI Foundation[32][45]
MCP, goose and AGENTS.md became founding projects; more than 10,000 MCP servers had been published.
confirmed
Google launches Gemini 3[42][43][44]
Gemini 3 Deep Think scored 45.1% on ARC-AGI-2; Google also introduced the Antigravity agentic development platform.
confirmed
Gemini Deep Think earns IMO gold-medal score[40][41]
The system solved five of six problems for 35 of 42 points, in natural language within the contest time.
confirmed
DeepSeek posts the DeepSeek-R1 paper[38][39]
It reported reasoning developed through pure reinforcement learning; the paper was later published in Nature.
confirmed
OpenAI publishes the o1 system card[37][1]
The card describes models trained with large-scale reinforcement learning to reason using chain of thought.
confirmed
o3 preview scores 75.7% on ARC-AGI-1[22]
A configuration using about 172 times more compute reached 87.5%, showing how much results depend on test-time compute.
confirmed
Anthropic introduces the Model Context Protocol[36]
confirmed
Anthropic releases computer use in public beta[2][35]
Claude 3.5 Sonnet could operate a computer through screenshots, cursor and keyboard, scoring 14.9% on OSWorld screenshot-only.
confirmed
OpenAI releases o1-preview and o1-mini[34]
The first models trained with reinforcement learning to reason in a chain of thought before answering.
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [2]
On 22 October 2024 Anthropic released computer use in public beta, letting Claude operate a computer by looking at a screen, moving a cursor, clicking and typing, and described the feature as experimental and error-prone. confirmedas of 2024-10-22
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku · Anthropic · 2024-10-22 (retrieved 2026-10-10)
- [3]
Anthropic said it would not make Claude Mythos Preview generally available until it had made progress on cybersecurity safeguards that detect and block the model's most dangerous outputs. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [4]
OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability, meaning that with tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [5]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [6]
On 12 June 2026, three days after Claude Fable 5's release, Anthropic said a US government export control directive required it to suspend all access to Fable 5 and Mythos 5 by foreign nationals, and that it had to disable both models for all customers to ensure compliance. confirmedas of 2026-06-12
- Statement on the US government directive to suspend access to Fable 5 and Mythos 5 · Anthropic · 2026-06-12 (retrieved 2026-10-10)
- Anthropic Releases and Temporarily Suspends Claude Fable 5 · InfoQ · Suspension section (retrieved 2026-10-10)
- [7]
Anthropic said the export controls on Fable 5 and Mythos 5 were lifted on 30 June 2026; Fable 5 returned to users globally on 1 July, and Mythos 5 access was restored for a set of US organisations after US government approval on 26 June. confirmedas of 2026-06-30
- Redeploying Fable 5 · Anthropic · 2026-06-30 (retrieved 2026-10-10)
- [8]
OpenAI's system card for GPT-6 Astra, dated 3 September 2026, calls it the most capable model OpenAI has ever broadly deployed and says it reasons through an extended chain of thought trained with reinforcement learning. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [9]
Anthropic released Claude Opus 5.5 on 22 September 2026, saying it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 (retrieved 2026-10-10)
- [10]
On September 30, 2026 Google DeepMind introduced Gemini 4 Argon, a model aimed at software engineering, enterprise knowledge work and cybersecurity defense. confirmedas of 2026-09-30
- Gemini 4 Argon: our next era of frontier intelligence · Google · 2026-09-30 · Post by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect (retrieved 2026-10-10)
- Google releases Gemini 4 Argon, called its most powerful model yet · TechCrunch · 2026-09-30 (retrieved 2026-10-10)
- [11]
On 29 September 2026 OpenAI introduced GPT-6.1 Sol, which it says has capabilities comparable to GPT-6 Astra and treats as Critical in cybersecurity and High in biological and chemical capability. confirmedas of 2026-09-29
- Addendum to GPT-6 Astra System Card: GPT-6.1 Sol · OpenAI · 2026-09-29 (retrieved 2026-10-10)
- [12]
Anthropic followed Claude Opus 5.5 with Claude Sonnet 5.5 on 28 September 2026 and Claude Haiku 5.5 on 7 October 2026. confirmedas of 2026-10-10
- Anthropic Newsroom · Anthropic · News listing, September and October 2026 (retrieved 2026-10-10)
- [13]
In its September 2026 roundup Google announced Gemini 4 Argon, a reasoning model with a 1-million-token output limit, rolling out first to cybersecurity professionals through its Fairwind Program. confirmedas of 2026-10-02
- The latest AI news we announced in September 2026 · Google · 2026-10-02 (retrieved 2026-10-10)
- [14]
Anthropic released Claude Mythos 5.1 on 1 September 2026 as its newest Mythos-class model, with access limited to a small set of vetted organisations. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [15]
Anthropic says Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1, with added safeguards for cybersecurity and biology. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [16]
On 10 September 2026 DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model, and said it was phasing out V4-Pro because tests put V4.1-Flash ahead of it. confirmedas of 2026-09-10
- DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient · DeepSeek · 2026-09-10 · The same post lists a "552B-parameter MoE" (retrieved 2026-10-10)
- [17]
The US Center for AI Standards and Innovation (CAISI) evaluated DeepSeek V4 Pro in April 2026 and estimated that its capabilities lagged the US frontier by about eight months, performing on CAISI's tests similarly to GPT-5 even though DeepSeek's self-reported results compared it with newer models. confirmedas of 2026-05-01
- CAISI Evaluation of DeepSeek V4 Pro · NIST · 2026-05-01 (retrieved 2026-10-10)
- [18]
On 6 October 2026 Anthropic expanded its Cyber Verification Program into three access tiers for security professionals, with access to models including Claude Mythos 5.1, and moved existing Project Glasswing members into its most permissive tier. confirmedas of 2026-10-06
- Expanding the Cyber Verification Program · Anthropic · 2026-10-06 (retrieved 2026-10-10)
- [19]
AISI estimated in February 2026 that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024, faster than its November 2025 estimate of 8 months, and said Claude Mythos Preview and GPT-5.5 then exceeded both trends. confirmedas of 2026-05-13
- How fast is autonomous AI cyber capability advancing? · UK AI Security Institute · 2026-05-13 (retrieved 2026-10-10)
- [20]
AISI reported in July 2026 that fixed-budget evaluations can systematically underestimate frontier agents' capabilities, because the test-time compute an agent may spend is a major driver of its measured capability. confirmedas of 2026-07-02
- More compute, more capability: Why AI agent evaluations need to account for test-time compute · UK AI Security Institute · 2026-07-02 (retrieved 2026-10-10)
- [21]
AISI found that the compute a task demands rises with how long it would take a skilled human, so the longest and hardest tasks are the first to be cut off by fixed budgets. confirmedas of 2026-07-02
- More compute, more capability: Why AI agent evaluations need to account for test-time compute · UK AI Security Institute · 2026-07-02 (retrieved 2026-10-10)
- [22]
In December 2024 the ARC Prize Foundation reported that a preview of OpenAI's o3 scored 75.7% on the ARC-AGI-1 semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute. confirmedas of 2024-12-20
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub · ARC Prize Foundation · 2024-12-20 (retrieved 2026-10-10)
- [23]
METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10
- METR says it can barely measure Claude Mythos · The Decoder · 2026-05-10 (retrieved 2026-10-10)
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates log, 8 May 2026 entry: Added Claude Mythos Preview (early); the point estimate appears in the page's chart data (retrieved 2026-10-10)
- [24]
METR's public time-horizons page carried a notice, in its 8 September 2026 update, that it is no longer actively updated. confirmedas of 2026-09-08
- Task-Completion Time Horizons of Frontier AI Models · METR (retrieved 2026-10-10)
- [25]
ARC Prize 2026 offers more than $2 million in prizes across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Prize details (retrieved 2026-10-10)
- [26]
As of October 2026 Anthropic describes Claude Fable 5.1 as a Mythos-level model for long-running projects, with safeguards that route many flagged cybersecurity and biology queries to less capable models. confirmedas of 2026-10-10
- Claude Fable · Anthropic (retrieved 2026-10-10)
- [27]
Anthropic said in June 2026 that it expected many other AI companies to have Mythos-class models within 6 to 12 months, possibly released without safeguards against misuse. confirmedas of 2026-06-02
- Expanding Project Glasswing · Anthropic · 2026-06-02 (retrieved 2026-10-10)
- [28]
METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [29]
OpenAI said its safeguards for GPT-6 Astra include universal monitoring of full trajectories, including chains of thought, across tool-using inference in its external deployment. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [30]
A 2025 paper by more than 40 researchers argued that monitoring reasoning models' chains of thought for intent to misbehave is a valuable safety opportunity, but that chain-of-thought monitorability may be fragile. confirmedas of 2025-12-07
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety · arXiv (Korbak et al.) · 2025-07-15 · Abstract (retrieved 2026-10-10)
- [31]
In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [32]
On 9 December 2025 the Linux Foundation formed the Agentic AI Foundation (AAIF), with Anthropic's MCP, Block's goose and OpenAI's AGENTS.md as founding projects. confirmedas of 2025-12-09
- Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF) · Linux Foundation · 2025-12-09 · Press release headline and first paragraph (retrieved 2026-10-10)
- [33]
The 2026 AI Index reports that US and Chinese models have traded places at the top of performance rankings several times since early 2025, and that as of March 2026 Anthropic's top model led by just 2.7%. confirmedas of 2026-03-31
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [34]
OpenAI released its first reasoning models, o1-preview and o1-mini, on 12 September 2024; the ARC Prize Foundation, writing on 13 September, said it had gained access to the newly released models over the previous 24 hours. confirmedas of 2024-09-13
- OpenAI o1 Results on ARC-AGI-Pub · ARC Prize Foundation · 2024-09-13 · Post dated 13 Sep 2024 (retrieved 2026-10-10)
- [35]
At launch Anthropic reported that Claude 3.5 Sonnet scored 14.9% on OSWorld in screenshot-only mode, against 7.8% for the next-best system. confirmedas of 2024-10-22
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku · Anthropic · 2024-10-22 · OSWorld results (retrieved 2026-10-10)
- [36]
Anthropic introduced the Model Context Protocol (MCP) on 25 November 2024 as an open standard for connecting AI assistants to the systems where data lives, such as content repositories, business tools and development environments. confirmedas of 2024-11-25
- Introducing the Model Context Protocol · Anthropic · 2024-11-25 (retrieved 2026-10-10)
- [37]
OpenAI's o1 system card was first posted on arXiv on 21 December 2024. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Submission history (v1, 21 December 2024) (retrieved 2026-10-10)
- [38]
The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Submission history and journal reference (retrieved 2026-10-10)
- [39]
DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [40]
In July 2025 an advanced version of Google DeepMind's Gemini with Deep Think achieved an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind (retrieved 2026-10-10)
- [41]
Google DeepMind said the 2025 IMO system worked end to end in natural language within the 4.5-hour contest limit, whereas its 2024 silver-medal system needed problems translated into the formal language Lean and two to three days of computation. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Main announcement (retrieved 2026-10-10)
- [42]
Google launched Gemini 3 on 18 November 2025, reporting 37.5% on Humanity's Last Exam without tools and 91.9% on GPQA Diamond for Gemini 3 Pro. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Benchmark section (retrieved 2026-10-10)
- [43]
Google reported in November 2025 that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2 and 41.0% on Humanity's Last Exam. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Deep Think section (retrieved 2026-10-10)
- [44]
Alongside Gemini 3, Google introduced Google Antigravity, an agentic development platform. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Product integration section (retrieved 2026-10-10)
- [45]
At the Agentic AI Foundation's launch in December 2025, the Linux Foundation said more than 10,000 MCP servers had been published. confirmedas of 2025-12-09
- Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF) · Linux Foundation · 2025-12-09 · MCP project description (retrieved 2026-10-10)
- [46]
Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [47]
The International AI Safety Report 2026, chaired by Yoshua Bengio, was published in February 2026 with contributions from more than 100 AI experts and an advisory panel spanning 29 nations, the UN, OECD and EU. confirmedas of 2026-02-24
- International AI Safety Report 2026 · arXiv (Bengio et al.) · 2026-02-24 · Abstract and author list (retrieved 2026-10-10)
- [48]
The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [49]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [50]
On 7 April 2026 Anthropic announced Project Glasswing, giving defenders of critical software access to Claude Mythos Preview, which it described as a general-purpose, unreleased frontier model. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [51]
Anthropic said Claude Mythos Preview found thousands of high-severity vulnerabilities, including some in every major operating system and web browser. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [52]
DeepSeek announced V4 preview models on 24 April 2026, V4-Pro with 1.6 trillion total and 49 billion active parameters and V4-Flash with 284 billion total and 13 billion active, both with a 1-million-token context. confirmedas of 2026-04-24
- DeepSeek V4 Preview Release · DeepSeek · 2026-04-24 (retrieved 2026-10-10)
- [53]
DeepSeek V4 offers both a thinking and a non-thinking mode, and its weights were released openly on Hugging Face. confirmedas of 2026-04-24
- DeepSeek V4 Preview Release · DeepSeek · 2026-04-24 (retrieved 2026-10-10)
- [54]
METR's time-horizons page lists an early version of Claude Mythos Preview as added to its measurements on 8 May 2026. confirmedas of 2026-05-08
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates section (retrieved 2026-10-10)
- [55]
On 2 June 2026 Anthropic extended Project Glasswing to about 150 more organisations in more than 15 countries, saying its roughly 50 initial partners had found more than 10,000 high- or critical-severity security flaws (company-reported). confirmedas of 2026-06-02
- Expanding Project Glasswing · Anthropic · 2026-06-02 (retrieved 2026-10-10)
- [56]
Anthropic released Claude Fable 5 and Claude Mythos 5 on 9 June 2026; they share one underlying model, with Mythos 5 carrying fewer safeguards and offered only to a small number of Project Glasswing partners for defensive cybersecurity. confirmedas of 2026-06-30
- Redeploying Fable 5 · Anthropic · 2026-06-30 (retrieved 2026-10-10)
- [57]
Anthropic said the US export control directive of 12 June 2026 followed a report in which Amazon researchers found a way to bypass Fable 5's safeguards and get it to identify software vulnerabilities. confirmedas of 2026-06-30
- Redeploying Fable 5 · Anthropic · 2026-06-30 (retrieved 2026-10-10)
- [58]
OpenAI's GPT-5.6, released with a system card on 9 July 2026, is a family of three models, Sol, Terra and Luna, which OpenAI treats as High capability in cybersecurity and in biological and chemical risk. confirmedas of 2026-07-09
- GPT-5.6 System Card · OpenAI · 2026-07-09 (retrieved 2026-10-10)
- [59]
OpenAI reported that GPT-5.6 Sol and Terra could find vulnerabilities and pieces of exploits but could not carry out autonomous end-to-end attacks against hardened targets in its testing. confirmedas of 2026-07-09
- GPT-5.6 System Card · OpenAI · 2026-07-09 (retrieved 2026-10-10)
- [60]
On 13 August 2026 DeepSeek made V4-Pro generally available, with selectable low, high and max reasoning effort. confirmedas of 2026-08-13
- DeepSeek-V4-Pro GA Release · DeepSeek · 2026-08-13 (retrieved 2026-10-10)
- [61]
On September 1, 2026 Anthropic released Claude Fable 5.1 and, for Project Glasswing participants, Claude Mythos 5.1, both priced at $10 per million input tokens and $50 per million output tokens. confirmedas of 2026-09-01
- Claude Platform release notes · Anthropic · September 1, 2026 entry (retrieved 2026-10-10)
- [62]
Anthropic reported that Project Glasswing partners uncovered at least 129,000 verified software vulnerabilities between April and July 2026, a figure it called a likely undercount based on partial survey data (company-reported). confirmedas of 2026-10-06
- Expanding the Cyber Verification Program · Anthropic · 2026-10-06 (retrieved 2026-10-10)
- [63]
On 8 October 2026 Anthropic launched the Anthropic Cyber Mission, starting with a Critical Infrastructure Defense Program for operational technology and free security scans of open-source projects by its strongest models. confirmedas of 2026-10-08
- Introducing the Anthropic Cyber Mission · Anthropic · 2026-10-08 (retrieved 2026-10-10)
- [64]
The ARC Prize Foundation reported that, with the provider-adapted harness, GPT-6 Astra used fewer actions than the human baseline on 96.0% of ARC-AGI-3 levels. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [65]
Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 · Benchmark table (retrieved 2026-10-10)
- [66]
Google also released Gemini 3.8 Flash and a Flash Cyber variant aimed at autonomous vulnerability discovery and automated code patching in September 2026. confirmedas of 2026-10-02
- The latest AI news we announced in September 2026 · Google · 2026-10-02 (retrieved 2026-10-10)
- [67]
OpenAI's 7 October 2026 system card says GPT-6 Sol and GPT-6 Luna replace GPT-5.6 models in ChatGPT and are rated High, but below Critical, in cybersecurity and biological and chemical capability. confirmedas of 2026-10-07
- GPT-6 Sol and GPT-6 Luna: October 2026 update · OpenAI Deployment Safety Hub · 2026-10-07 (retrieved 2026-10-10)
- [68]
OpenAI said neither GPT-6 Sol nor GPT-6 Luna reaches its High threshold for AI self-improvement. confirmedas of 2026-10-07
- GPT-6 Sol and GPT-6 Luna: October 2026 update · OpenAI Deployment Safety Hub · 2026-10-07 (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Frontier AI tracker: reasoning models and agents." ContentLora, updated Oct 10, 2026. https://contentlora.com/events/frontier-ai-tracker
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.