Frontier AI
Reasoning models, AI agents and the frontier of large-scale AI systems: what they can do and how they are built.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.Updated
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.Updated
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.Updated
- WikiMETRMETR is an AI evaluation group known for measuring how long a task AI agents can complete, and for trials of AI's effect on developers.Updated
- WikiModel Context Protocol (MCP)The Model Context Protocol is an open standard for connecting AI apps and agents to tools and data, now run by the Linux Foundation.Updated
- WikiOpenAI o1OpenAI o1 is a reasoning model series trained with reinforcement learning to think in a chain of thought before answering.Updated
- WikiSWE-benchSWE-bench tests whether AI can fix real GitHub issues. How it works, its Verified subset, and how scores rose from 2% to near 100%.Updated
- WikiTest-time computeTest-time compute is the computing power an AI model spends while answering. Spending more of it is how reasoning models improve.Updated
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.Updated
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.Updated
- ExplainerHow AI agents workWhat an AI agent is and how it works: language models using tools in a loop, computer use, MCP connectors and long tasks.Updated
- ExplainerHow frontier AI capabilities are measuredHow researchers measure what frontier AI can do: benchmarks like SWE-bench, Humanity's Last Exam and ARC-AGI, and METR time horizons.Updated
- ExplainerHow large language models workA plain guide to large language models: the Transformer, scaling laws and human-feedback training, at beginner and expert level.Updated
- ExplainerHow reasoning models workHow AI reasoning models think step by step: chain of thought, reinforcement learning on checkable tasks and test-time compute.Updated