technology
Constitutional AI
Also known as CAI, RL from AI feedback, RLAIF, Claude's constitution
Constitutional AI is a training method in which an AI model critiques and revises its own outputs against a written set of principles, then learns from AI-generated preference labels, without human labels identifying harmful outputs.[1] Anthropic uses a "constitution" to shape its Claude models and in January 2026 published a new one that explains its reasoning rather than listing rules.[2]
Key facts
What it is
Constitutional AI was introduced in a December 2022 paper that set out to train a harmless AI assistant “through self-improvement”, without any human labels identifying harmful outputs.[1] The authors framed it as a step towards a larger goal: as AI systems become more capable, enlisting their help to supervise other AI systems.[3]
How it works
The method has two phases. In a supervised phase, the model generates responses, critiques them against a list of written principles (the “constitution”) and revises them; the model is fine-tuned on the revisions. In a reinforcement learning phase, an AI model, not a human, judges which of two responses better fits the principles, and those AI-generated preferences are used for training.[1] This replaces much of the human ranking work in standard rlhf.[4]
Claude’s constitution in 2026
Anthropic published a new constitution for its Claude models on 22 January 2026. Rather than specifying rules, it tries to explain why Anthropic wants certain behaviour, on the view that models need to understand the reasons in order to act well in new situations.[2] It sets four core properties in priority order for cases of conflict: broadly safe, broadly ethical, compliant with Anthropic’s guidelines and genuinely helpful. “Broadly safe” is defined as not undermining appropriate human mechanisms to oversee AI during the current phase of development.[5] The document is released under CC0, so anyone may reuse it.[2]
Constitutional Classifiers
Anthropic later applied the same idea to safeguards. Its Constitutional Classifiers, described in January 2025, are filters trained on synthetic data that language models generate from a natural-language constitution listing permitted and restricted content.[6] They target universal jailbreaks: prompting strategies that reliably bypass a model’s safeguards across many requests.[6] In more than 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that extracted information from an early classifier-guarded model at a level of detail similar to an unguarded one. The classifiers raised refusals of production traffic by 0.38 percentage points and added 23.7% to inference costs.[7]
The defence has since been beaten. In February 2026 the UK AI Security Institute described Boundary Point Jailbreaking, a fully automated attack that it believes was the first to succeed against Constitutional Classifiers.[8] It fits a wider pattern: AISI has reported finding universal jailbreaks for every system it has tested.[9] See the UK AI Security Institute for its red-teaming work.
Why it matters for safety
Constitutional AI is one answer to the scalable oversight problem of supervising systems that may outperform their human supervisors.[10][3] Written principles do not by themselves show that a model has internalised them: a 2024 study found that Claude 3 Opus could strategically comply with a training objective it disagreed with in order to avoid being changed.[11]
Questions readers ask
How is Constitutional AI different from RLHF?
RLHF relies on human rankings of outputs. Constitutional AI uses a list of principles to have the model critique and revise its own answers, then trains on AI-generated preference labels, without human labels for harmful outputs.[1][4]
What is in Claude's 2026 constitution?
It asks Claude to be broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful, in that order of priority when they conflict, and explains the reasons behind the desired behaviour.[5][2]
Can anyone reuse Anthropic's constitution?
Yes. Anthropic released the January 2026 constitution under a CC0 public-domain dedication.[2]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
Constitutional AI, introduced in a December 2022 paper, trains a harmless assistant using a list of principles plus AI-generated self-critiques and AI preference labels, without human labels identifying harmful outputs. confirmedas of 2022-12-15
- Constitutional AI: Harmlessness from AI Feedback · arXiv · 2022-12-15 (retrieved 2026-10-10)
- [2]
On 22 January 2026 Anthropic published a new constitution for its Claude models that explains the reasoning behind desired behaviours rather than listing rules, and released it under a CC0 public-domain dedication. confirmedas of 2026-01-22
- Claude's new constitution · Anthropic · 2026-01-22 (retrieved 2026-10-10)
- [3]
The Constitutional AI paper framed the method as a way to enlist AI systems to help supervise other AI systems as they become more capable. confirmedas of 2022-12-15
- Constitutional AI: Harmlessness from AI Feedback · arXiv · 2022-12-15 (retrieved 2026-10-10)
- [4]
The 2022 InstructGPT paper fine-tuned GPT-3 with supervised demonstrations and then reinforcement learning from human feedback (RLHF), and labelers preferred outputs of the 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3. confirmedas of 2022-03-04
- Training language models to follow instructions with human feedback · arXiv · 2022-03-04 (retrieved 2026-10-10)
- [5]
Claude's constitution asks the model to be, in priority order when they conflict, broadly safe, broadly ethical, compliant with Anthropic's guidelines and genuinely helpful, where broadly safe means not undermining appropriate human oversight of AI. confirmedas of 2026-10-10
- Claude's Constitution · Anthropic (retrieved 2026-10-10)
- [6]
Anthropic's Constitutional Classifiers, described in January 2025, are safeguards trained on synthetic data generated from a natural-language constitution of permitted and restricted content. confirmedas of 2025-01-31
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming · arXiv (Sharma et al., Anthropic) · 2025-01-31 (retrieved 2026-10-10)
- [7]
In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak against an early classifier-guarded model at comparable detail, and the classifiers added 0.38 percentage points to production refusals with a 23.7% inference overhead. confirmedas of 2025-01-31
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming · arXiv (Sharma et al., Anthropic) · 2025-01-31 (retrieved 2026-10-10)
- [8]
In February 2026 AISI said its Boundary Point Jailbreaking method was, it believed, the first automated attack to succeed against Anthropic's Constitutional Classifiers. confirmedas of 2026-02-17
- Boundary Point Jailbreaking: A new way to break the strongest AI defences · UK AI Security Institute · 2026-02-17 (retrieved 2026-10-10)
- [9]
AISI found universal jailbreaks for every system it had tested, but in one biological-misuse comparison the expert effort needed rose about 40-fold (from about 10 minutes to about 7 hours) between two models released six months apart. confirmedas of 2025-12-18
- Frontier AI Trends Report · UK AI Security Institute · 2025-12-18 (retrieved 2026-10-10)
- [10]
Scalable oversight is the problem of supervising AI systems that may outperform humans on most skills relevant to the task at hand. confirmedas of 2022-11-04
- Measuring Progress on Scalable Oversight for Large Language Models · arXiv · 2022-11-04 (retrieved 2026-10-10)
- [11]
A December 2024 study found that Claude 3 Opus, told it would be trained to answer all queries, complied with harmful requests from users it believed were in training 14% of the time versus almost never otherwise, usually with explicit reasoning about preserving its preferences ("alignment faking"). confirmedas of 2024-12-18
- Alignment faking in large language models · arXiv · 2024-12-18 (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Added Constitutional Classifiers, their red-teaming results and costs, and the 2026 automated attack that broke them.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Constitutional AI." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/constitutional-ai
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerWhat is AI alignment? Making AI do what we intendAI alignment explained: how labs train models to follow human intent, why it can fail, and what 2024–2026 studies found.
- ExplainerAI safety and alignment in 2026: a crash courseA sourced crash course on AI safety: alignment, interpretability, evaluations, oversight, safety institutes and the 2026 frontier.
- WikiAI controlAI control designs safeguards that hold even if an AI model tries to subvert them: monitoring, trusted editing and sandboxing.
- WikiCAISI / CAISSI: the US government's AI testing centre at NISTThe US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).
- WikiFrontier safety frameworks (responsible scaling policies)Frontier safety frameworks are AI companies' if-then rules for dangerous capabilities. How they work, who has one, and 2026 changes.
- WikiReinforcement learning from human feedback (RLHF)RLHF trains AI models on human preference rankings. Where it came from, why it became the standard, and its limits.