concept
Frontier safety frameworks (responsible scaling policies)
Also known as responsible scaling policies, RSP, Frontier Safety Framework, Preparedness Framework, frontier AI safety policies
Frontier safety frameworks, such as Anthropic's Responsible Scaling Policy and Google DeepMind's Frontier Safety Framework, are company policies that set capability thresholds and the safeguards, or pauses, that must follow when a model crosses them.[1] Twelve companies had published one by December 2025,[2] and California's SB 53 and the EU's code of practice now expect frontier developers to publish such frameworks.[3][4]
Key facts
What they are
A frontier safety framework is an “if-then” commitment: if testing shows a model has reached a dangerous capability, then specific protections must be in place before it is deployed or trained further. METR’s survey of these policies finds common elements: capability thresholds, model weight security, deployment mitigations, conditions for halting deployment or development, full capability elicitation in evaluations, and accountability.[1]
Origins and spread
Anthropic introduced its Responsible Scaling Policy in September 2023.[5] At the AI Seoul Summit in May 2024, 20 companies signed Frontier AI Safety Commitments to publish frameworks focused on severe risks,[6] including a pledge not to develop or deploy a model at all if mitigations cannot keep risks below their thresholds.[7] By December 2025 METR counted twelve companies with published policies, including Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI and NVIDIA.[2]
Major frameworks in 2026
- Anthropic RSP. Version 3.0 took effect on 24 February 2026 and was revised four times by July 2026 (version 3.4, 8 July).[5] Version 3.0 added published Frontier Safety Roadmaps, Risk Reports covering deployed models, and clarified AI Safety Level (ASL) standards.[8]
- Google DeepMind Frontier Safety Framework. Updated in September 2025 and to version 3.1 in April 2026, it added a critical capability level for harmful manipulation, protocols for misalignment risks, and lower “tracked capability levels” to flag risks earlier.[9]
- OpenAI Preparedness Framework. In August 2026 OpenAI paused reinforcement learning on its newest models for two weeks and kept its largest planned training run on hold,[10] after preliminary evidence that an upcoming model might reach the framework’s critical cybersecurity level.[11]
Frameworks in action
2026 showed how these thresholds work in practice. Anthropic revised its RSP in April, May and July, putting out versions 3.1 to 3.4 within about five months.[12] OpenAI treated its GPT-5.6 models, released in July, as High but not Critical for cybersecurity and for biological and chemical risk.[13] It then rated GPT-6 Astra at the Critical cybersecurity level, meaning that with tools and access it can find unknown flaws and build new exploits across many well-protected systems.[14] GPT-6.1 Sol, released at the end of September, is also treated as Critical for cybersecurity and uses the same safeguards as Astra.[15] At Anthropic, the most capable Mythos models remain limited to vetted organisations, while the same model is sold with extra safeguards as Claude Fable 5.1.[16][17]
Frameworks are also extending to training. According to MIXED, OpenAI published draft guidelines in September 2026 saying a structured safety case, covering alignment training, containment and monitoring, should be required before continuing any frontier reinforcement learning run, with senior leaders each able to veto it.[18] And Anthropic began turning its pledge of outside scrutiny into practice: in September it announced embedded evaluation with Accenture, with each firm expecting to invest at least $1 billion over five years.[19] It acknowledged that no standards yet define what such evaluators may see or how they should report.[20]
From voluntary to required
Frameworks began as voluntary pledges but are now referenced in law and regulation. California’s SB 53, signed on 29 September 2025, requires large frontier developers to publish a framework, report critical safety incidents and protect whistleblowers.[3] The EU’s General-Purpose AI Code of Practice, published on 10 July 2025, includes a safety and security chapter for providers of models with systemic risk.[4]
Limits
Frameworks depend on evaluations that can detect dangerous capabilities, and the International AI Safety Report 2026 says reliable pre-deployment testing has become harder because models can tell tests from real use.[21] In September 2026 Anthropic’s chief executive argued that companies should go further and deliberately slow capability gains, with embedded third-party evaluators.[22][23]
Questions readers ask
What is a responsible scaling policy?
A company policy that defines dangerous capability thresholds and commits to stronger safeguards, or to halting deployment or development, if a model crosses one. Anthropic's version has existed since September 2023.[1][5]
Are these frameworks legally required?
Increasingly. California's SB 53 requires large frontier developers to publish a safety framework, and the EU's code of practice sets safety and security practices for providers of the most advanced models.[3][4]
Has any company actually paused because of its framework?
In August 2026 OpenAI paused reinforcement learning on its newest models for two weeks after preliminary evidence that an upcoming model might reach the "critical" cyber level of its Preparedness Framework.[10][11]
What changed in Anthropic's 2026 policy?
Version 3.0, effective 24 February 2026, added published Frontier Safety Roadmaps and Risk Reports and clarified its AI Safety Level standards.[8]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
METR identified common elements of frontier safety policies, including capability thresholds, model weight security, deployment mitigations, conditions for halting deployment or development, full capability elicitation in evaluations, and accountability. confirmedas of 2025-12-16
- Common Elements of Frontier AI Safety Policies · METR · 2025-12-16 (retrieved 2026-10-10)
- [2]
As of December 2025, METR counted twelve companies with published frontier AI safety policies, including Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI and NVIDIA. confirmedas of 2025-12-16
- Common Elements of Frontier AI Safety Policies · METR · 2025-12-16 · Introduction (retrieved 2026-10-10)
- [3]
California's SB 53, the Transparency in Frontier Artificial Intelligence Act, signed on 29 September 2025, requires large frontier developers to publish a safety framework, creates a channel for reporting critical safety incidents to the state Office of Emergency Services, and protects whistleblowers. confirmedas of 2025-09-29
- Governor Newsom signs SB 53, advancing California's world-leading artificial intelligence industry · Office of the Governor of California · 2025-09-29 (retrieved 2026-10-10)
- [4]
The European Commission published the General-Purpose AI Code of Practice on 10 July 2025, with chapters on transparency, copyright, and safety and security; the safety and security chapter applies only to providers of general-purpose models with systemic risk. confirmedas of 2025-07-10
- The General-Purpose AI Code of Practice · European Commission (retrieved 2026-10-10)
- [5]
Anthropic's Responsible Scaling Policy was first introduced in September 2023; version 3.0 took effect on 24 February 2026 and version 3.4 on 8 July 2026. confirmedas of 2026-10-10
- Anthropic's Responsible Scaling Policy · Anthropic · Version history (retrieved 2026-10-10)
- [6]
At the AI Seoul Summit in May 2024, 20 companies signed Frontier AI Safety Commitments to publish safety frameworks focused on severe risks, including thresholds beyond which risks are unacceptable. confirmedas of 2025-02-07
- Frontier AI Safety Commitments, AI Seoul Summit 2024 · GOV.UK (Department for Science, Innovation and Technology) (retrieved 2026-10-10)
- [7]
Under the Seoul commitments, signatories committed not to develop or deploy a model at all if mitigations cannot keep risks below their thresholds. confirmedas of 2025-02-07
- Frontier AI Safety Commitments, AI Seoul Summit 2024 · GOV.UK (Department for Science, Innovation and Technology) (retrieved 2026-10-10)
- [8]
Version 3.0 of Anthropic's Responsible Scaling Policy introduced published Frontier Safety Roadmaps, Risk Reports assessing risks across deployed models, and clarified AI Safety Level (ASL) security and deployment standards. confirmedas of 2026-02-24
- Anthropic's Responsible Scaling Policy · Anthropic · Version 3.0 summary (retrieved 2026-10-10)
- [9]
Google DeepMind's Frontier Safety Framework, updated in September 2025 and again to version 3.1 in April 2026, added a critical capability level for harmful manipulation, protocols for misalignment risks such as interference with operator control, and lower "tracked capability levels" to spot risks sooner. confirmedas of 2026-04-17
- Strengthening our Frontier Safety Framework · Google DeepMind · 2025-09-22 (retrieved 2026-10-10)
- [10]
In August 2026 OpenAI paused reinforcement learning training on its latest models intended for deployment for two weeks, said its largest planned frontier RL run remained on hold pending more evidence of alignment, and expanded the coverage of its monitoring systems. confirmedas of 2026-08-19
- OpenAI puts major frontier AI training run on hold over cyber risks · Help Net Security · 2026-08-19 (retrieved 2026-10-10)
- OpenAI: We'll hit pause on model reinforcement learning for safety · Constellation Research · 2026-08-18 (retrieved 2026-10-10)
- [11]
Help Net Security and DataBreachToday reported that the pause followed preliminary evidence about the cybersecurity capabilities of OpenAI's upcoming Astra model, which Help Net Security said may meet the Critical cybersecurity threshold of OpenAI's Preparedness Framework; Constellation Research likewise reported that OpenAI had noted Astra may have critical cyber capabilities. confirmedas of 2026-08-19
- OpenAI puts major frontier AI training run on hold over cyber risks · Help Net Security · 2026-08-19 (retrieved 2026-10-10)
- OpenAI Pauses Frontier Model Training for Safety Review · DataBreachToday · 2026-08-19 (retrieved 2026-10-10)
- OpenAI: We'll hit pause on model reinforcement learning for safety · Constellation Research · 2026-08-18 (retrieved 2026-10-10)
- [12]
Anthropic revised its Responsible Scaling Policy four times in 2026 after version 3.0, with versions 3.1 (2 April), 3.2 (29 April), 3.3 (26 May) and 3.4 (8 July). confirmedas of 2026-10-10
- Anthropic's Responsible Scaling Policy · Anthropic (retrieved 2026-10-10)
- [13]
OpenAI's GPT-5.6, released with a system card on 9 July 2026, is a family of three models, Sol, Terra and Luna, which OpenAI treats as High capability in cybersecurity and in biological and chemical risk. confirmedas of 2026-07-09
- GPT-5.6 System Card · OpenAI · 2026-07-09 (retrieved 2026-10-10)
- [14]
OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability, meaning that with tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [15]
On 29 September 2026 OpenAI introduced GPT-6.1 Sol, which it says has capabilities comparable to GPT-6 Astra and treats as Critical in cybersecurity and High in biological and chemical capability. confirmedas of 2026-09-29
- Addendum to GPT-6 Astra System Card: GPT-6.1 Sol · OpenAI · 2026-09-29 (retrieved 2026-10-10)
- [16]
Anthropic released Claude Mythos 5.1 on 1 September 2026 as its newest Mythos-class model, with access limited to a small set of vetted organisations. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [17]
Anthropic says Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1, with added safeguards for cybersecurity and biology. confirmedas of 2026-10-10
- Claude Mythos · Anthropic (retrieved 2026-10-10)
- [18]
MIXED reported that on 28 September 2026 OpenAI published draft guidelines saying structured safety cases, covering alignment training, containment and monitoring, should be required before continuing any frontier reinforcement learning run, and that senior leaders should each be able to veto a run. reportedas of 2026-10-04
- OpenAI says senior leaders should each be able to veto a frontier training run · MIXED (The Decoder) · 2026-10-04 (retrieved 2026-10-10)
- [19]
On 18 September 2026 Anthropic announced a non-exclusive partnership with Accenture, led by its Faculty unit, to embed independent evaluators inside Anthropic, with both companies expecting to invest at least $1 billion each in this capacity over five years. confirmedas of 2026-09-18
- Partnering with Accenture on embedded evaluation · Anthropic · 2026-09-18 (retrieved 2026-10-10)
- [20]
Anthropic said there are as yet no standards for what information embedded evaluators should access or how they should report findings. confirmedas of 2026-09-18
- Partnering with Accenture on embedded evaluation · Anthropic · 2026-09-18 (retrieved 2026-10-10)
- [21]
The International AI Safety Report 2026 says reliable pre-deployment safety testing has become harder because models increasingly distinguish between test settings and real deployment. confirmedas of 2026-02-03
- International AI Safety Report 2026 · International AI Safety Report · 2026-02-03 (retrieved 2026-10-10)
- [22]
In a September 2026 essay, Anthropic CEO Dario Amodei argued that the industry must slow the pace of AI capability improvements, citing AI's growing ability to build the next generation of AI and the OpenAI–Hugging Face incident. confirmedas of 2026-09-12
- We Must Pace the Frontier · Dario Amodei (personal essay, Anthropic CEO) · 2026-09-12 (retrieved 2026-10-10)
- [23]
In the same essay Amodei said Anthropic was unilaterally committing to embed third-party evaluators with access comparable to internal risk assessors and rights to publish findings, and proposed coordination among democratic countries and with authoritarian governments. confirmedas of 2026-09-12
- We Must Pace the Frontier · Dario Amodei (personal essay, Anthropic CEO) · 2026-09-12 (retrieved 2026-10-10)
- [24]
The UK AI Safety Institute was set up following the first AI Safety Summit on 1–2 November 2023 and evaluated 16 models in its first year, working with companies including OpenAI, Google DeepMind and Anthropic. confirmedas of 2024-11-01
- Our first year · UK AI Security Institute (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Frontier safety frameworks (responsible scaling policies)." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/frontier-safety-frameworks
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow AI safety testing works: evals, red teams and thresholdsHow frontier AI models are tested before release: dangerous-capability evals, jailbreak red-teaming, and why testing got harder.
- ExplainerAI safety and alignment in 2026: a crash courseA sourced crash course on AI safety: alignment, interpretability, evaluations, oversight, safety institutes and the 2026 frontier.
- AnalysisCan AI safety keep pace with AI capabilities?The central debate in AI safety in 2026: are evaluations, interpretability and oversight keeping up with fast-rising capabilities?
- WikiCAISI / CAISSI: the US government's AI testing centre at NISTThe US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).
- WikiUK AI Security Institute (AISI)The UK AI Security Institute tests frontier AI models for national-security risks. Its history, tools and key findings.
- DevelopingAI safety tracker: alignment, interpretability and evals in 2026Live tracker of AI safety milestones: interpretability results, evaluations, safety frameworks, incidents and institutes.