{"event":"frontier-ai-tracker","title":"Frontier AI tracker: reasoning models and agents","state":"developing","updated":"2026-10-10T00:00:00.000Z","entries":[{"id":"18yy5tc","at":"2026-10-08T00:00:00.000Z","title":"Anthropic launches its Cyber Mission","status":"confirmed","cites":[{"id":"frontier-ai-cyber-mission","n":63,"url":"https://www.anthropic.com/news/anthropic-cyber-mission","publisher":"Anthropic","short":"Anthropic"}]},{"id":"3x3iu9","at":"2026-10-07T00:00:00.000Z","title":"Anthropic releases Claude Haiku 5.5","status":"confirmed","cites":[{"id":"frontier-ai-anthropic-5-5-family","n":12,"url":"https://www.anthropic.com/news","publisher":"Anthropic","short":"Anthropic"}]},{"id":"dimfv5","at":"2026-10-07T00:00:00.000Z","title":"OpenAI ships GPT-6 Sol and Luna in ChatGPT","status":"confirmed","body":"Both are rated High, below Critical, for cyber and bio capability; neither reaches High for AI self-improvement.","cites":[{"id":"frontier-ai-gpt-6-sol-luna","n":67,"url":"https://deploymentsafety.openai.com/gpt-6-october","publisher":"OpenAI Deployment Safety Hub","short":"OpenAI Deployment Safety Hub"},{"id":"frontier-ai-gpt-6-sol-luna-self-improvement","n":68,"url":"https://deploymentsafety.openai.com/gpt-6-october","publisher":"OpenAI Deployment Safety Hub","short":"OpenAI Deployment Safety Hub"}]},{"id":"1lrkhjx","at":"2026-10-06T00:00:00.000Z","title":"Anthropic folds Glasswing into an expanded Cyber Verification Program","status":"confirmed","body":"Anthropic reported (company figures) that Glasswing partners found at least 129,000 verified vulnerabilities from April to July.","cites":[{"id":"frontier-ai-cvp-expansion","n":18,"url":"https://www.anthropic.com/news/cyber-verification-program","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-cvp-glasswing-vulns","n":62,"url":"https://www.anthropic.com/news/cyber-verification-program","publisher":"Anthropic","short":"Anthropic"}]},{"id":"1sfor14","at":"2026-09-30T00:00:00.000Z","title":"Google DeepMind introduces Gemini 4 Argon, first for cyber defenders","status":"confirmed","cites":[{"id":"google-deepmind-gemini-4-argon","n":10,"url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/","publisher":"Google","short":"Google"},{"id":"frontier-ai-gemini-4-argon","n":13,"url":"https://blog.google/innovation-and-ai/technology/ai/google-ai-updates-september-2026/","publisher":"Google","short":"Google"},{"id":"frontier-ai-gemini-flash-cyber","n":66,"url":"https://blog.google/innovation-and-ai/technology/ai/google-ai-updates-september-2026/","publisher":"Google","short":"Google"}]},{"id":"1ri8g4h","at":"2026-09-29T00:00:00.000Z","title":"OpenAI's GPT-6.1 Sol rated Critical for cyber","status":"confirmed","cites":[{"id":"frontier-ai-o1x-gpt-6-1-sol","n":11,"url":"https://deploymentsafety.openai.com/gpt-6-1-sol","publisher":"OpenAI","short":"OpenAI"}]},{"id":"236z9t","at":"2026-09-28T00:00:00.000Z","title":"Anthropic releases Claude Sonnet 5.5","status":"confirmed","cites":[{"id":"frontier-ai-anthropic-5-5-family","n":12,"url":"https://www.anthropic.com/news","publisher":"Anthropic","short":"Anthropic"}]},{"id":"1szzh4r","at":"2026-09-22T00:00:00.000Z","title":"Anthropic releases Claude Opus 5.5","status":"confirmed","body":"Anthropic says it performs at Claude Fable 5.1 level on most work at 40% lower cost than Opus 5.","cites":[{"id":"frontier-ai-opus-5-5-launch","n":9,"url":"https://www.anthropic.com/news/claude-opus-5-5","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-opus-5-5-benchmarks","n":65,"url":"https://www.anthropic.com/news/claude-opus-5-5","publisher":"Anthropic","short":"Anthropic"}]},{"id":"1uj1elg","at":"2026-09-10T00:00:00.000Z","title":"DeepSeek releases V4.1-Flash and phases out V4-Pro","status":"confirmed","cites":[{"id":"frontier-ai-r1-v4-1-flash","n":16,"url":"https://api-docs.deepseek.com/news/news260910","publisher":"DeepSeek","short":"DeepSeek"}]},{"id":"13jv39l","at":"2026-09-08T00:00:00.000Z","title":"METR stops actively updating its time-horizons page","status":"confirmed","cites":[{"id":"frontier-ai-metr-page-paused","n":24,"url":"https://metr.org/time-horizons/","publisher":"METR","short":"METR"}]},{"id":"13mpz65","at":"2026-09-03T00:00:00.000Z","title":"OpenAI releases GPT-6 Astra, rated Critical for cyber","status":"confirmed","cites":[{"id":"frontier-ai-gpt-6-astra","n":8,"url":"https://deploymentsafety.openai.com/gpt-6-astra","publisher":"OpenAI Deployment Safety Hub","short":"OpenAI Deployment Safety Hub"},{"id":"frontier-ai-gpt-6-astra-critical-cyber","n":4,"url":"https://deploymentsafety.openai.com/gpt-6-astra","publisher":"OpenAI Deployment Safety Hub","short":"OpenAI Deployment Safety Hub"}]},{"id":"12grsjw","at":"2026-09-03T00:00:00.000Z","title":"GPT-6 Astra scores 62.7% on ARC-AGI-3","status":"confirmed","cites":[{"id":"frontier-ai-arc-agi-3-astra","n":5,"url":"https://arcprize.org/blog/astra","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"},{"id":"frontier-ai-arc-agi-3-astra-efficiency","n":64,"url":"https://arcprize.org/blog/astra","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"}]},{"id":"s9i5rm","at":"2026-09-01T00:00:00.000Z","title":"Anthropic releases Claude Fable 5.1 and Mythos 5.1","status":"confirmed","cites":[{"id":"anthropic-claude-5-1-models","n":61,"url":"https://platform.claude.com/docs/en/release-notes/overview","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-mythos-5-1","n":14,"url":"https://www.anthropic.com/claude/mythos","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-mythos-fable-same-model","n":15,"url":"https://www.anthropic.com/claude/mythos","publisher":"Anthropic","short":"Anthropic"}]},{"id":"47tbf8","at":"2026-08-13T00:00:00.000Z","title":"DeepSeek makes V4-Pro generally available","status":"confirmed","cites":[{"id":"frontier-ai-r1-v4-pro-ga","n":60,"url":"https://api-docs.deepseek.com/news/news260813","publisher":"DeepSeek","short":"DeepSeek"}]},{"id":"1rllonl","at":"2026-07-09T00:00:00.000Z","title":"OpenAI publishes the GPT-5.6 system card","status":"confirmed","body":"OpenAI treated Sol, Terra and Luna as High, not Critical, for cyber and bio capability.","cites":[{"id":"frontier-ai-o1x-gpt-5-6-family","n":58,"url":"https://deploymentsafety.openai.com/gpt-5-6","publisher":"OpenAI","short":"OpenAI"},{"id":"frontier-ai-o1x-gpt-5-6-cyber-limit","n":59,"url":"https://deploymentsafety.openai.com/gpt-5-6","publisher":"OpenAI","short":"OpenAI"}]},{"id":"2hhk6","at":"2026-06-30T00:00:00.000Z","title":"US lifts export controls on Fable 5 and Mythos 5","status":"confirmed","body":"Fable 5 returned globally on 1 July; Mythos 5 was restored for a set of US organisations.","cites":[{"id":"frontier-ai-mythos-export-lifted","n":7,"url":"https://www.anthropic.com/news/redeploying-fable-5","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-mythos-directive-cause","n":57,"url":"https://www.anthropic.com/news/redeploying-fable-5","publisher":"Anthropic","short":"Anthropic"}]},{"id":"13wq9vw","at":"2026-06-12T00:00:00.000Z","title":"Anthropic pulls Claude Fable 5 under a US export control directive","status":"confirmed","body":"Three days after release, Anthropic said a US government export control directive required it to disable Fable 5 and Mythos 5 for all customers.","cites":[{"id":"frontier-ai-fable-5-suspended","n":6,"url":"https://www.anthropic.com/news/fable-mythos-access","publisher":"Anthropic","short":"Anthropic"}]},{"id":"132zvyf","at":"2026-06-09T00:00:00.000Z","title":"Anthropic releases Claude Fable 5 and Claude Mythos 5","status":"confirmed","body":"The two share one model; Mythos 5, with fewer safeguards, went only to Glasswing partners.","cites":[{"id":"frontier-ai-mythos-5-launch","n":56,"url":"https://www.anthropic.com/news/redeploying-fable-5","publisher":"Anthropic","short":"Anthropic"}]},{"id":"29ezby","at":"2026-06-02T00:00:00.000Z","title":"Anthropic extends Project Glasswing to about 150 more organisations","status":"confirmed","cites":[{"id":"frontier-ai-mythos-glasswing-expansion","n":55,"url":"https://www.anthropic.com/news/expanding-project-glasswing","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-mythos-copycat-warning","n":27,"url":"https://www.anthropic.com/news/expanding-project-glasswing","publisher":"Anthropic","short":"Anthropic"}]},{"id":"1dwthuj","at":"2026-05-08T00:00:00.000Z","title":"METR measures Claude Mythos Preview at 16+ hours","status":"confirmed","body":"METR said the estimate was at the upper end of what its task suite can measure.","cites":[{"id":"frontier-ai-tracker-metr-mythos-added","n":54,"url":"https://metr.org/time-horizons/","publisher":"METR","short":"METR"},{"id":"frontier-ai-metr-mythos-16h","n":23,"url":"https://the-decoder.com/metr-says-it-can-barely-measure-claude-mythos-palo-alto-networks-warns-of-autonomous-ai-attackers/","publisher":"The Decoder","short":"The Decoder"}]},{"id":"1h2gx26","at":"2026-04-24T00:00:00.000Z","title":"DeepSeek releases V4 preview with open weights","status":"confirmed","cites":[{"id":"frontier-ai-deepseek-v4-params","n":52,"url":"https://api-docs.deepseek.com/news/news260424","publisher":"DeepSeek","short":"DeepSeek"},{"id":"frontier-ai-deepseek-v4-modes","n":53,"url":"https://api-docs.deepseek.com/news/news260424","publisher":"DeepSeek","short":"DeepSeek"}]},{"id":"m7faz3","at":"2026-04-07T00:00:00.000Z","title":"Anthropic announces Claude Mythos Preview and Project Glasswing","status":"confirmed","body":"Anthropic withheld general release over cybersecurity risk and gave access to defenders of critical software.","cites":[{"id":"frontier-ai-glasswing-launch","n":50,"url":"https://www.anthropic.com/glasswing","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-glasswing-vulns","n":51,"url":"https://www.anthropic.com/glasswing","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-glasswing-not-ga","n":3,"url":"https://www.anthropic.com/glasswing","publisher":"Anthropic","short":"Anthropic"}]},{"id":"13wuho","at":"2026-03-25T00:00:00.000Z","title":"ARC-AGI-3 launches; frontier AI scores 0.51%","status":"confirmed","cites":[{"id":"frontier-ai-arc-agi-3-launch","n":49,"url":"https://arcprize.org/blog/arc-agi-3-launch","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"},{"id":"frontier-ai-arc-prize-2026","n":25,"url":"https://arcprize.org/blog/arc-agi-3-launch","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"}]},{"id":"1g3j8jl","at":"2026-02-24T00:00:00.000Z","title":"International AI Safety Report 2026 published","status":"confirmed","body":"The report highlights inference-time scaling and uneven, \"jagged\" capabilities.","cites":[{"id":"frontier-ai-intl-report-about","n":47,"url":"https://arxiv.org/abs/2602.21012","publisher":"arXiv (Bengio et al.)","short":"arXiv"},{"id":"frontier-ai-intl-report-inference-scaling","n":48,"url":"https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026","publisher":"International AI Safety Report (UK DSIT 2026/001)","short":"International AI Safety…"}]},{"id":"1tjfqpx","at":"2026-01-29T00:00:00.000Z","title":"METR releases Time Horizon 1.1","status":"confirmed","body":"METR estimated a doubling time of about 89 days since 2024, with Claude Opus 4.5 at about 320 minutes.","cites":[{"id":"frontier-ai-metr-th11","n":31,"url":"https://metr.org/blog/2026-1-29-time-horizon-1-1/","publisher":"METR","short":"METR"},{"id":"frontier-ai-metr-th11-models","n":46,"url":"https://metr.org/blog/2026-1-29-time-horizon-1-1/","publisher":"METR","short":"METR"}]},{"id":"uolaaj","at":"2025-12-09T00:00:00.000Z","title":"Linux Foundation forms the Agentic AI Foundation","status":"confirmed","body":"MCP, goose and AGENTS.md became founding projects; more than 10,000 MCP servers had been published.","cites":[{"id":"frontier-ai-aaif-formation","n":32,"url":"https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation","publisher":"Linux Foundation","short":"Linux Foundation"},{"id":"frontier-ai-mcp-servers-count","n":45,"url":"https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation","publisher":"Linux Foundation","short":"Linux Foundation"}]},{"id":"rfim2z","at":"2025-11-18T00:00:00.000Z","title":"Google launches Gemini 3","status":"confirmed","body":"Gemini 3 Deep Think scored 45.1% on ARC-AGI-2; Google also introduced the Antigravity agentic development platform.","cites":[{"id":"frontier-ai-gemini-3-launch","n":42,"url":"https://blog.google/products/gemini/gemini-3/","publisher":"Google","short":"Google"},{"id":"frontier-ai-gemini-3-arc-agi-2","n":43,"url":"https://blog.google/products/gemini/gemini-3/","publisher":"Google","short":"Google"},{"id":"frontier-ai-gemini-3-antigravity","n":44,"url":"https://blog.google/products/gemini/gemini-3/","publisher":"Google","short":"Google"}]},{"id":"1qcqe43","at":"2025-07-21T00:00:00.000Z","title":"Gemini Deep Think earns IMO gold-medal score","status":"confirmed","body":"The system solved five of six problems for 35 of 42 points, in natural language within the contest time.","cites":[{"id":"frontier-ai-imo-gold","n":40,"url":"https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/","publisher":"Google DeepMind","short":"Google DeepMind"},{"id":"frontier-ai-imo-natural-language","n":41,"url":"https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/","publisher":"Google DeepMind","short":"Google DeepMind"}]},{"id":"d7t48n","at":"2025-01-22T00:00:00.000Z","title":"DeepSeek posts the DeepSeek-R1 paper","status":"confirmed","body":"It reported reasoning developed through pure reinforcement learning; the paper was later published in Nature.","cites":[{"id":"frontier-ai-deepseek-r1-published","n":38,"url":"https://arxiv.org/abs/2501.12948","publisher":"arXiv (DeepSeek-AI); published in Nature 645","short":"arXiv; published in Nature…"},{"id":"frontier-ai-deepseek-r1-pure-rl","n":39,"url":"https://arxiv.org/abs/2501.12948","publisher":"arXiv (DeepSeek-AI); published in Nature 645","short":"arXiv; published in Nature…"}]},{"id":"hoqjg4","at":"2024-12-21T00:00:00.000Z","title":"OpenAI publishes the o1 system card","status":"confirmed","body":"The card describes models trained with large-scale reinforcement learning to reason using chain of thought.","cites":[{"id":"frontier-ai-tracker-o1-system-card","n":37,"url":"https://arxiv.org/abs/2412.16720","publisher":"arXiv (OpenAI)","short":"arXiv"},{"id":"frontier-ai-o1-rl-cot","n":1,"url":"https://arxiv.org/abs/2412.16720","publisher":"arXiv (OpenAI)","short":"arXiv"}]},{"id":"19xi9k0","at":"2024-12-20T00:00:00.000Z","title":"o3 preview scores 75.7% on ARC-AGI-1","status":"confirmed","body":"A configuration using about 172 times more compute reached 87.5%, showing how much results depend on test-time compute.","cites":[{"id":"frontier-ai-arcagi-o3","n":22,"url":"https://arcprize.org/blog/oai-o3-pub-breakthrough","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"}]},{"id":"3qa8l9","at":"2024-11-25T00:00:00.000Z","title":"Anthropic introduces the Model Context Protocol","status":"confirmed","cites":[{"id":"frontier-ai-mcp-launch","n":36,"url":"https://www.anthropic.com/news/model-context-protocol","publisher":"Anthropic","short":"Anthropic"}]},{"id":"rmqg1e","at":"2024-10-22T00:00:00.000Z","title":"Anthropic releases computer use in public beta","status":"confirmed","body":"Claude 3.5 Sonnet could operate a computer through screenshots, cursor and keyboard, scoring 14.9% on OSWorld screenshot-only.","cites":[{"id":"frontier-ai-computer-use-launch","n":2,"url":"https://www.anthropic.com/news/3-5-models-and-computer-use","publisher":"Anthropic","short":"Anthropic"},{"id":"frontier-ai-computer-use-osworld","n":35,"url":"https://www.anthropic.com/news/3-5-models-and-computer-use","publisher":"Anthropic","short":"Anthropic"}]},{"id":"2x7cwr","at":"2024-09-12T00:00:00.000Z","title":"OpenAI releases o1-preview and o1-mini","status":"confirmed","body":"The first models trained with reinforcement learning to reason in a chain of thought before answering.","cites":[{"id":"frontier-ai-o1-preview-release","n":34,"url":"https://arcprize.org/blog/openai-o1-results-arc-prize","publisher":"ARC Prize Foundation","short":"ARC Prize Foundation"}]}]}