Explainer
Frontier AI in 2026: a crash course
Frontier AI in 2026 is defined by two shifts: reasoning models, trained with reinforcement learning to think step by step before answering,[1] and AI agents, which use tools in a loop to complete multi-step tasks.[2] By October 2026 the most capable models can work autonomously on software tasks that take experts many hours,[3] and labs have started gating them over cybersecurity risk.[4][5]
This crash course takes you from what a large language model is to where the frontier of AI stands in October 2026. Read the pages in the order listed in the course panel: fundamentals first, then key methods, models and benchmarks, then the open debates, and finally the live tracker. Most sections below have a beginner and an expert version. The field’s own assessments stress uncertainty: the International AI Safety Report 2026 says progress through 2030 is uncertain but that current trends are consistent with continued improvement.[6]
Why frontier AI matters
The most advanced AI systems now do work that used to need skilled people. They fix real software bugs, operate ordinary computers and solve olympiad mathematics.[7][8][9] Global corporate AI investment reached $581.7 billion in 2025.[10] The same abilities raise risks: Anthropic said its Claude Mythos Preview model found thousands of serious security flaws, including some in every major operating system and web browser.[11]
Two capability trends define the field. Agent task horizons, as measured by METR, doubled about every seven months from 2019 to early 2025 and about every 89 days since 2024 under its revised method.[12][13] Meanwhile, inference-time scaling has become a main lever for capability.[14] Dual-use risk is now driving deployment decisions: OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability,[5] and Anthropic withheld general release of Mythos Preview pending safeguards.[4] Transparency is falling, with average Foundation Model Transparency Index scores down to 40 from 58.[15]
The map of the field
Think of four layers. Foundation models are large language models trained on huge amounts of text (how they work).[16] Reasoning models are trained to think step by step before answering.[1] Agents use reasoning models plus tools, such as a code runner or web browser, to finish tasks.[17] Evaluation is the science of measuring all this, through benchmarks and studies.[18]
- Pre-training and scaling: Transformers trained on next-token prediction, with loss falling as a power law in parameters, data and compute.[19][20]
- Post-training: RLHF, then reinforcement learning on verifiable tasks that produces long chains of thought.[21][22]
- Inference-time scaling: spending more compute per answer, sequentially or in parallel (test-time-compute).[23][24]
- Agents and tooling: tool-use loops, computer use and connector standards such as the model-context-protocol.[2][25][26]
- Evaluation: swe-bench, OSWorld, Humanity’s Last Exam, arc-agi and METR time horizons.[27][28][29]
Key ideas in one place
- Chain of thought: the model writes out intermediate steps, which improves complex reasoning.[30]
- Reinforcement learning: the model practises on checkable problems and is rewarded for correct answers; reasoning behaviours such as self-checking can emerge.[31]
- Agent loop: act, look at the result, decide the next step.[2]
- Jagged frontier: models can excel at hard tasks yet fail simple ones, like reading an analog clock.[32][33]
OpenAI’s openai-o1 series is trained with large-scale RL to reason using chain of thought.[1] DeepSeek’s deepseek-r1 reported that pure RL without human-labelled trajectories suffices and was published in Nature.[22][34] DeepSeek reported its gains on verifiable domains such as mathematics and coding.[35] A key open issue is faithfulness: in Anthropic tests, models mentioned hints they used only 25% to 39% of the time.[36]
Who the main players are
A handful of labs build the most capable models. OpenAI makes the GPT-6 family,[37] Anthropic makes Claude, including the Mythos class,[38] and Google DeepMind makes Gemini.[39] DeepSeek releases strong models with openly downloadable weights.[40] Independent groups such as METR and the ARC Prize Foundation test what the models can do.[13][41] The chips they run on come from firms such as NVIDIA, which designs its processors but has them made by foundries such as TSMC; NVIDIA’s data-center revenue rose 68% to $193.7 billion in fiscal 2026.[42][43] See how chips are made and AI chip supply concentration.
- OpenAI: GPT-6 Astra (September 2026) and GPT-6 Sol and Luna (October 2026), with system cards on its deployment safety hub.[37][44][45]
- Anthropic: gated Claude Mythos models, now Mythos 5.1 for vetted organisations; Claude Fable 5.1, the same model with safeguards; and the Claude 5.5 family from September 2026.[46][47][48]
- Google DeepMind: Gemini 3 (November 2025), IMO gold with Deep Think, and Gemini 4 Argon (September 2026).[49][9][39]
- DeepSeek: R1, then V4-Pro with 1.6 trillion total parameters and open weights.[50][40]
- Standards and evaluators: the Linux Foundation’s Agentic AI Foundation (MCP), METR and the ARC Prize Foundation.[51][13][52]
The 2026 AI Index reports that US and Chinese models have traded the top spot several times since early 2025.[53]
Where the frontier is in October 2026
In early 2025, the best AI could finish software tasks that take a person about 50 minutes, half the time.[12] By March 2026, an early Claude Mythos model reached at least 16 hours.[3] On ARC-AGI-3, a set of puzzle games people solve easily, AI went from 0.51% in March 2026 to 62.7% in September.[41][54] But evidence that these tools make professionals faster is still weak.[55]
Anthropic reports 93.9% on SWE-bench Verified for Mythos Preview[56] and, for Opus 5.5, 81.8% on OSWorld 2.1 and 67.7% on Humanity’s Last Exam with tools.[57] GPT-6 Astra’s ARC-AGI-3 result rises from 62.7% to 99.9% with a provider-adapted harness.[54] Measurement limits matter: METR calls estimates above 16 hours unreliable and has stopped actively updating its public time-horizons page.[3][58] The International AI Safety Report 2026 judges progress through 2030 uncertain but consistent with continued improvement.[6]
Open debates
Two questions run through the field: whether reasoning models’ abilities are general and scalable or brittle, and whether rapid benchmark gains for agents reflect real-world usefulness.[59][60] Each has its own analysis page in this course, and the tracker logs new milestones as they are confirmed.
Questions readers ask
What is a reasoning model?
A language model trained with large-scale reinforcement learning to work through a chain of thought before answering, using more computing power at answer time.[1][14]
What is an AI agent?
A system in which a language model dynamically directs its own process and tool use to complete a task, typically by acting, checking results and repeating in a loop.[17][2]
How long can today's AI agents work on their own?
METR estimated a 50% time horizon of at least 16 hours for an early version of Claude Mythos Preview in March 2026, at the edge of what it can measure.[3]
Is the US still ahead of China in AI models?
Only narrowly. The 2026 AI Index reports that US and Chinese models have traded places at the top several times since early 2025, with Anthropic's top model leading by 2.7% in March 2026.[53]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [2]
Anthropic describes agents as typically language models using tools based on feedback from their environment in a loop, checking results such as tool outputs or code execution at each step. confirmedas of 2024-12-19
- Building effective agents · Anthropic · 2024-12-19 (retrieved 2026-10-10)
- [3]
METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10
- METR says it can barely measure Claude Mythos · The Decoder · 2026-05-10 (retrieved 2026-10-10)
- Task-Completion Time Horizons of Frontier AI Models · METR · Updates log, 8 May 2026 entry: Added Claude Mythos Preview (early); the point estimate appears in the page's chart data (retrieved 2026-10-10)
- [4]
Anthropic said it would not make Claude Mythos Preview generally available until it had made progress on cybersecurity safeguards that detect and block the model's most dangerous outputs. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [5]
OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability, meaning that with tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [6]
The International AI Safety Report 2026 concludes that the trajectory of AI progress through 2030 is uncertain but that current trends are consistent with continued improvement. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [7]
The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [8]
The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [9]
In July 2025 an advanced version of Google DeepMind's Gemini with Deep Think achieved an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind (retrieved 2026-10-10)
- [10]
The 2026 AI Index reports that global corporate AI investment reached $581.7 billion in 2025, up 130% from the year before. confirmedas of 2026-04-01
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
- [11]
Anthropic said Claude Mythos Preview found thousands of high-severity vulnerabilities, including some in every major operating system and web browser. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [12]
METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [13]
In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [14]
The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [15]
The 2026 AI Index reports that average scores on the Foundation Model Transparency Index fell to 40 from 58 the year before, with the most capable models disclosing the least. confirmedas of 2026-04-01
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
- [16]
The 2020 GPT-3 paper reported that scaling up language models greatly improves few-shot performance, with the 175-billion-parameter model doing new tasks from text prompts alone, without gradient updates. confirmedas of 2020-05-28
- Language Models are Few-Shot Learners · arXiv (Brown et al.) · 2020-05-28 · Abstract (retrieved 2026-10-10)
- [17]
Anthropic defines agents as systems in which language models dynamically direct their own processes and tool use, keeping control over how they accomplish a task. confirmedas of 2024-12-19
- Building effective agents · Anthropic · 2024-12-19 (retrieved 2026-10-10)
- [18]
METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [19]
The Transformer architecture, introduced in a 2017 paper, is based solely on attention mechanisms and dispenses with recurrence and convolutions. confirmedas of 2017-06-12
- Attention Is All You Need · arXiv (Vaswani et al.) · 2017-06-12 · Abstract (retrieved 2026-10-10)
- [20]
A 2020 study found that language-model loss falls as a power law as model size, dataset size and training compute grow, with some trends spanning more than seven orders of magnitude. confirmedas of 2020-01-23
- Scaling Laws for Neural Language Models · arXiv (Kaplan et al.) · 2020-01-23 · Abstract (retrieved 2026-10-10)
- [21]
InstructGPT was trained in two stages, supervised fine-tuning on human-written demonstrations and then reinforcement learning from human feedback (RLHF) using human rankings of model outputs. confirmedas of 2022-03-04
- Training language models to follow instructions with human feedback · arXiv (Ouyang et al.) · 2022-03-04 · Abstract (retrieved 2026-10-10)
- [22]
DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [23]
A 2024 study of two mechanisms, search against process-based verifier models and adaptive revision of responses, found that allocating test-time compute adaptively per prompt improved test-time compute efficiency by more than 4x over a best-of-N baseline, and that in some settings a smaller model given extra inference compute could outperform a 14x larger model. confirmedas of 2024-08-06
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters · arXiv (Snell et al.) · 2024-08-06 · Abstract (retrieved 2026-10-10)
- [24]
Google DeepMind described Deep Think as exploring multiple solution paths in parallel and said the IMO model was trained with reinforcement learning techniques focused on multi-step problem solving. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Technical approach section (retrieved 2026-10-10)
- [25]
On 22 October 2024 Anthropic released computer use in public beta, letting Claude operate a computer by looking at a screen, moving a cursor, clicking and typing, and described the feature as experimental and error-prone. confirmedas of 2024-10-22
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku · Anthropic · 2024-10-22 (retrieved 2026-10-10)
- [26]
Anthropic introduced the Model Context Protocol (MCP) on 25 November 2024 as an open standard for connecting AI assistants to the systems where data lives, such as content repositories, business tools and development environments. confirmedas of 2024-11-25
- Introducing the Model Context Protocol · Anthropic · 2024-11-25 (retrieved 2026-10-10)
- [27]
SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (the arXiv page renders the counts in LaTeX as $2,294$ and $12$) (retrieved 2026-10-10)
- SWE-bench Lite · SWE-bench (retrieved 2026-10-10)
- [28]
Humanity's Last Exam, released in January 2025, is a benchmark of 2,500 expert-written questions across dozens of subjects, created because models were scoring over 90% on popular benchmarks such as MMLU. confirmedas of 2025-01-24
- Humanity's Last Exam · arXiv (Phan et al.) · 2025-01-24 · Abstract (retrieved 2026-10-10)
- Humanity's Last Exam · Center for AI Safety and Scale AI (retrieved 2026-10-10)
- [29]
The ARC Prize Foundation describes ARC-AGI-3 as an interactive reasoning benchmark in which AI agents must explore novel environments, acquire goals on the fly, build adaptable world models and learn continuously. confirmedas of 2026-10-10
- ARC-AGI-3 · ARC Prize Foundation (retrieved 2026-10-10)
- [30]
A 2022 study showed that prompting large language models to generate a chain of thought, a series of intermediate reasoning steps, significantly improves their performance on complex reasoning. confirmedas of 2022-01-28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models · arXiv (Wei et al.) · 2022-01-28 · Abstract (retrieved 2026-10-10)
- [31]
According to the DeepSeek-R1 paper, reinforcement learning led the model to develop behaviours such as self-reflection, verification and dynamic strategy adaptation. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [32]
The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [33]
The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time, an example of uneven capabilities. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [34]
The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Submission history and journal reference (retrieved 2026-10-10)
- [35]
The DeepSeek-R1 paper reports that its reinforcement-learning-trained models did better than conventionally supervised models on verifiable tasks such as mathematics, coding competitions and STEM problems. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [36]
Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03
- Reasoning models don't always say what they think · Anthropic (retrieved 2026-10-10)
- [37]
OpenAI's system card for GPT-6 Astra, dated 3 September 2026, calls it the most capable model OpenAI has ever broadly deployed and says it reasons through an extended chain of thought trained with reinforcement learning. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [38]
On 7 April 2026 Anthropic announced Project Glasswing, giving defenders of critical software access to Claude Mythos Preview, which it described as a general-purpose, unreleased frontier model. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [39]
In its September 2026 roundup Google announced Gemini 4 Argon, a reasoning model with a 1-million-token output limit, rolling out first to cybersecurity professionals through its Fairwind Program. confirmedas of 2026-10-02
- The latest AI news we announced in September 2026 · Google · 2026-10-02 (retrieved 2026-10-10)
- [40]
DeepSeek V4 offers both a thinking and a non-thinking mode, and its weights were released openly on Hugging Face. confirmedas of 2026-04-24
- DeepSeek V4 Preview Release · DeepSeek · 2026-04-24 (retrieved 2026-10-10)
- [41]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [42]
NVIDIA does not manufacture its own chips; it uses foundries such as TSMC and Samsung Electronics. confirmedas of 2026-02-25
- NVIDIA Corporation Form 10-K for fiscal year ended January 25, 2026 · NVIDIA (SEC filing) · 2026-02-25 · Item 1. Business, Manufacturing (retrieved 2026-10-10)
- [43]
NVIDIA's Data Center revenue for fiscal year 2026 rose 68% to $193.7 billion. confirmedas of 2026-01-25
- NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2026 · NVIDIA Newsroom · 2026-02-25 · Data Center (retrieved 2026-10-10)
- [44]
OpenAI's 7 October 2026 system card says GPT-6 Sol and GPT-6 Luna replace GPT-5.6 models in ChatGPT and are rated High, but below Critical, in cybersecurity and biological and chemical capability. confirmedas of 2026-10-07
- GPT-6 Sol and GPT-6 Luna: October 2026 update · OpenAI Deployment Safety Hub · 2026-10-07 (retrieved 2026-10-10)
- [45]
OpenAI's deployment safety hub lists system cards for GPT-5.6 (9 July 2026), GPT-6 Astra (3 September 2026), a GPT-6.1 Sol addendum (29 September 2026) and GPT-6 Sol and Luna (7 October 2026). confirmedas of 2026-10-10
- OpenAI Deployment Safety Hub · OpenAI · System card list (retrieved 2026-10-10)
- [46]
Anthropic released Claude Mythos 5.1 on 1 September 2026 as its newest Mythos-class model, with access limited to a small set of vetted organisations. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [47]
Anthropic says Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1, with added safeguards for cybersecurity and biology. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [48]
Anthropic released Claude Opus 5.5 on 22 September 2026, saying it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 (retrieved 2026-10-10)
- [49]
Google launched Gemini 3 on 18 November 2025, reporting 37.5% on Humanity's Last Exam without tools and 91.9% on GPQA Diamond for Gemini 3 Pro. confirmedas of 2025-11-18
- A new era of intelligence with Gemini 3 · Google · 2025-11-18 · Benchmark section (retrieved 2026-10-10)
- [50]
DeepSeek announced V4 preview models on 24 April 2026, V4-Pro with 1.6 trillion total and 49 billion active parameters and V4-Flash with 284 billion total and 13 billion active, both with a 1-million-token context. confirmedas of 2026-04-24
- DeepSeek V4 Preview Release · DeepSeek · 2026-04-24 (retrieved 2026-10-10)
- [51]
On 9 December 2025 the Linux Foundation formed the Agentic AI Foundation (AAIF), with Anthropic's MCP, Block's goose and OpenAI's AGENTS.md as founding projects. confirmedas of 2025-12-09
- Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF) · Linux Foundation · 2025-12-09 · Press release headline and first paragraph (retrieved 2026-10-10)
- [52]
ARC Prize 2026 offers more than $2 million in prizes across an ARC-AGI-3 agent competition on Kaggle and an ARC-AGI-2 Grand Prize for the best open-source solution. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Prize details (retrieved 2026-10-10)
- [53]
The 2026 AI Index reports that US and Chinese models have traded places at the top of performance rankings several times since early 2025, and that as of March 2026 Anthropic's top model led by just 2.7%. confirmedas of 2026-03-31
- Inside the AI Index: 12 Takeaways from the 2026 Report · Stanford HAI (retrieved 2026-10-10)
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [54]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [55]
METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [56]
Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 · Benchmark table (retrieved 2026-10-10)
- [57]
Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 · Benchmark table (retrieved 2026-10-10)
- [58]
METR's public time-horizons page carried a notice, in its 8 September 2026 update, that it is no longer actively updated. confirmedas of 2026-09-08
- Task-Completion Time Horizons of Frontier AI Models · METR (retrieved 2026-10-10)
- [59]
A 2025 Apple study using controllable puzzles found that large reasoning models face a complete accuracy collapse beyond certain problem complexities. confirmedas of 2025-11-20
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · arXiv (Shojaee et al., Apple) · 2025-06-07 · Abstract (retrieved 2026-10-10)
- [60]
In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · arXiv (METR) · Abstract (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Linked the lab, chipmaker and supply-chain pages and updated Anthropic's models to Mythos 5.1.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Frontier AI in 2026: a crash course." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/frontier-ai
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- DevelopingFrontier AI tracker: reasoning models and agentsA dated timeline of frontier AI milestones in reasoning models and AI agents, from o1 and DeepSeek-R1 to GPT-6, Claude Mythos and Gemini 4.
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.
- WikiModel Context Protocol (MCP)The Model Context Protocol is an open standard for connecting AI apps and agents to tools and data, now run by the Linux Foundation.
- WikiSWE-benchSWE-bench tests whether AI can fix real GitHub issues. How it works, its Verified subset, and how scores rose from 2% to near 100%.
- WikiTest-time computeTest-time compute is the computing power an AI model spends while answering. Spending more of it is how reasoning models improve.