Skip to content
ContentLora

    Tip: press / anywhere to search.

    concept

    Frontier safety frameworks (responsible scaling policies)

    Also known as responsible scaling policies, RSP, Frontier Safety Framework, Preparedness Framework, frontier AI safety policies

    Frontier safety frameworks, such as Anthropic's Responsible Scaling Policy and Google DeepMind's Frontier Safety Framework, are company policies that set capability thresholds and the safeguards, or pauses, that must follow when a model crosses them.[1] Twelve companies had published one by December 2025,[2] and California's SB 53 and the EU's code of practice now expect frontier developers to publish such frameworks.[3][4]

    Editor reviewedUpdated AI safety and alignmentArtificial intelligenceTech policy
    Key facts

    What they are

    A frontier safety framework is an “if-then” commitment: if testing shows a model has reached a dangerous capability, then specific protections must be in place before it is deployed or trained further. METR’s survey of these policies finds common elements: capability thresholds, model weight security, deployment mitigations, conditions for halting deployment or development, full capability elicitation in evaluations, and accountability.[1]

    Origins and spread

    Anthropic introduced its Responsible Scaling Policy in September 2023.[5] At the AI Seoul Summit in May 2024, 20 companies signed Frontier AI Safety Commitments to publish frameworks focused on severe risks,[6] including a pledge not to develop or deploy a model at all if mitigations cannot keep risks below their thresholds.[7] By December 2025 METR counted twelve companies with published policies, including Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI and NVIDIA.[2]

    Major frameworks in 2026

    • Anthropic RSP. Version 3.0 took effect on 24 February 2026 and was revised four times by July 2026 (version 3.4, 8 July).[5] Version 3.0 added published Frontier Safety Roadmaps, Risk Reports covering deployed models, and clarified AI Safety Level (ASL) standards.[8]
    • Google DeepMind Frontier Safety Framework. Updated in September 2025 and to version 3.1 in April 2026, it added a critical capability level for harmful manipulation, protocols for misalignment risks, and lower “tracked capability levels” to flag risks earlier.[9]
    • OpenAI Preparedness Framework. In August 2026 OpenAI paused reinforcement learning on its newest models for two weeks and kept its largest planned training run on hold,[10] after preliminary evidence that an upcoming model might reach the framework’s critical cybersecurity level.[11]

    Frameworks in action

    2026 showed how these thresholds work in practice. Anthropic revised its RSP in April, May and July, putting out versions 3.1 to 3.4 within about five months.[12] OpenAI treated its GPT-5.6 models, released in July, as High but not Critical for cybersecurity and for biological and chemical risk.[13] It then rated GPT-6 Astra at the Critical cybersecurity level, meaning that with tools and access it can find unknown flaws and build new exploits across many well-protected systems.[14] GPT-6.1 Sol, released at the end of September, is also treated as Critical for cybersecurity and uses the same safeguards as Astra.[15] At Anthropic, the most capable Mythos models remain limited to vetted organisations, while the same model is sold with extra safeguards as Claude Fable 5.1.[16][17]

    Frameworks are also extending to training. According to MIXED, OpenAI published draft guidelines in September 2026 saying a structured safety case, covering alignment training, containment and monitoring, should be required before continuing any frontier reinforcement learning run, with senior leaders each able to veto it.[18] And Anthropic began turning its pledge of outside scrutiny into practice: in September it announced embedded evaluation with Accenture, with each firm expecting to invest at least $1 billion over five years.[19] It acknowledged that no standards yet define what such evaluators may see or how they should report.[20]

    From voluntary to required

    Frameworks began as voluntary pledges but are now referenced in law and regulation. California’s SB 53, signed on 29 September 2025, requires large frontier developers to publish a framework, report critical safety incidents and protect whistleblowers.[3] The EU’s General-Purpose AI Code of Practice, published on 10 July 2025, includes a safety and security chapter for providers of models with systemic risk.[4]

    Limits

    Frameworks depend on evaluations that can detect dangerous capabilities, and the International AI Safety Report 2026 says reliable pre-deployment testing has become harder because models can tell tests from real use.[21] In September 2026 Anthropic’s chief executive argued that companies should go further and deliberately slow capability gains, with embedded third-party evaluators.[22][23]

    Questions readers ask

    What is a responsible scaling policy?

    A company policy that defines dangerous capability thresholds and commits to stronger safeguards, or to halting deployment or development, if a model crosses one. Anthropic's version has existed since September 2023.[1][5]

    Are these frameworks legally required?

    Increasingly. California's SB 53 requires large frontier developers to publish a safety framework, and the EU's code of practice sets safety and security practices for providers of the most advanced models.[3][4]

    Has any company actually paused because of its framework?

    In August 2026 OpenAI paused reinforcement learning on its newest models for two weeks after preliminary evidence that an upcoming model might reach the "critical" cyber level of its Preparedness Framework.[10][11]

    What changed in Anthropic's 2026 policy?

    Version 3.0, effective 24 February 2026, added published Frontier Safety Roadmaps and Risk Reports and clarified its AI Safety Level standards.[8]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      METR identified common elements of frontier safety policies, including capability thresholds, model weight security, deployment mitigations, conditions for halting deployment or development, full capability elicitation in evaluations, and accountability. confirmedas of 2025-12-16

    2. [2]

      As of December 2025, METR counted twelve companies with published frontier AI safety policies, including Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI and NVIDIA. confirmedas of 2025-12-16

    3. [3]

      California's SB 53, the Transparency in Frontier Artificial Intelligence Act, signed on 29 September 2025, requires large frontier developers to publish a safety framework, creates a channel for reporting critical safety incidents to the state Office of Emergency Services, and protects whistleblowers. confirmedas of 2025-09-29

    4. [4]

      The European Commission published the General-Purpose AI Code of Practice on 10 July 2025, with chapters on transparency, copyright, and safety and security; the safety and security chapter applies only to providers of general-purpose models with systemic risk. confirmedas of 2025-07-10

    5. [5]

      Anthropic's Responsible Scaling Policy was first introduced in September 2023; version 3.0 took effect on 24 February 2026 and version 3.4 on 8 July 2026. confirmedas of 2026-10-10

    6. [6]

      At the AI Seoul Summit in May 2024, 20 companies signed Frontier AI Safety Commitments to publish safety frameworks focused on severe risks, including thresholds beyond which risks are unacceptable. confirmedas of 2025-02-07

    7. [7]

      Under the Seoul commitments, signatories committed not to develop or deploy a model at all if mitigations cannot keep risks below their thresholds. confirmedas of 2025-02-07

    8. [8]

      Version 3.0 of Anthropic's Responsible Scaling Policy introduced published Frontier Safety Roadmaps, Risk Reports assessing risks across deployed models, and clarified AI Safety Level (ASL) security and deployment standards. confirmedas of 2026-02-24

    9. [9]

      Google DeepMind's Frontier Safety Framework, updated in September 2025 and again to version 3.1 in April 2026, added a critical capability level for harmful manipulation, protocols for misalignment risks such as interference with operator control, and lower "tracked capability levels" to spot risks sooner. confirmedas of 2026-04-17

    10. [10]

      In August 2026 OpenAI paused reinforcement learning training on its latest models intended for deployment for two weeks, said its largest planned frontier RL run remained on hold pending more evidence of alignment, and expanded the coverage of its monitoring systems. confirmedas of 2026-08-19

    11. [11]

      Help Net Security and DataBreachToday reported that the pause followed preliminary evidence about the cybersecurity capabilities of OpenAI's upcoming Astra model, which Help Net Security said may meet the Critical cybersecurity threshold of OpenAI's Preparedness Framework; Constellation Research likewise reported that OpenAI had noted Astra may have critical cyber capabilities. confirmedas of 2026-08-19

    12. [12]

      Anthropic revised its Responsible Scaling Policy four times in 2026 after version 3.0, with versions 3.1 (2 April), 3.2 (29 April), 3.3 (26 May) and 3.4 (8 July). confirmedas of 2026-10-10

    13. [13]

      OpenAI's GPT-5.6, released with a system card on 9 July 2026, is a family of three models, Sol, Terra and Luna, which OpenAI treats as High capability in cybersecurity and in biological and chemical risk. confirmedas of 2026-07-09

    14. [14]

      OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability, meaning that with tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems. confirmedas of 2026-09-03

    15. [15]

      On 29 September 2026 OpenAI introduced GPT-6.1 Sol, which it says has capabilities comparable to GPT-6 Astra and treats as Critical in cybersecurity and High in biological and chemical capability. confirmedas of 2026-09-29

    16. [16]

      Anthropic released Claude Mythos 5.1 on 1 September 2026 as its newest Mythos-class model, with access limited to a small set of vetted organisations. confirmedas of 2026-10-10

    17. [17]

      Anthropic says Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1, with added safeguards for cybersecurity and biology. confirmedas of 2026-10-10

    18. [18]

      MIXED reported that on 28 September 2026 OpenAI published draft guidelines saying structured safety cases, covering alignment training, containment and monitoring, should be required before continuing any frontier reinforcement learning run, and that senior leaders should each be able to veto a run. reportedas of 2026-10-04

    19. [19]

      On 18 September 2026 Anthropic announced a non-exclusive partnership with Accenture, led by its Faculty unit, to embed independent evaluators inside Anthropic, with both companies expecting to invest at least $1 billion each in this capacity over five years. confirmedas of 2026-09-18

    20. [20]

      Anthropic said there are as yet no standards for what information embedded evaluators should access or how they should report findings. confirmedas of 2026-09-18

    21. [21]

      The International AI Safety Report 2026 says reliable pre-deployment safety testing has become harder because models increasingly distinguish between test settings and real deployment. confirmedas of 2026-02-03

    22. [22]

      In a September 2026 essay, Anthropic CEO Dario Amodei argued that the industry must slow the pace of AI capability improvements, citing AI's growing ability to build the next generation of AI and the OpenAI–Hugging Face incident. confirmedas of 2026-09-12

    23. [23]

      In the same essay Amodei said Anthropic was unilaterally committing to embed third-party evaluators with access comparable to internal risk assessors and rights to publish findings, and proposed coordination among democratic countries and with authoritarian governments. confirmedas of 2026-09-12

    24. [24]

      The UK AI Safety Institute was set up following the first AI Safety Summit on 1–2 November 2023 and evaluated 16 models in its first year, working with companies including OpenAI, Google DeepMind and Anthropic. confirmedas of 2024-11-01

    Revision history (2)
    1. Page created.
    2. Added how OpenAI's Preparedness ratings were applied to GPT-5.6, GPT-6 Astra and GPT-6.1 Sol, Anthropic's 2026 RSP revisions, OpenAI's draft safety-case guidelines (reported) and the Anthropic-Accenture embedded evaluation deal.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Frontier safety frameworks (responsible scaling policies)." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/frontier-safety-frameworks

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.