AI safety and alignment
Research on making AI systems reliable, interpretable and aligned with human intent, and how it is evaluated.
- WikiAI controlAI control designs safeguards that hold even if an AI model tries to subvert them: monitoring, trusted editing and sandboxing.Updated
- WikiCAISI / CAISSI: the US government's AI testing centre at NISTThe US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).Updated
- WikiConstitutional AIConstitutional AI trains models with written principles and AI feedback instead of only human labels. How it works and what changed in 2026.Updated
- WikiFrontier safety frameworks (responsible scaling policies)Frontier safety frameworks are AI companies' if-then rules for dangerous capabilities. How they work, who has one, and 2026 changes.Updated
- WikiReinforcement learning from human feedback (RLHF)RLHF trains AI models on human preference rankings. Where it came from, why it became the standard, and its limits.Updated
- WikiScalable oversightScalable oversight asks how humans can supervise AI that outperforms them. Debate, AI-assisted judging and weak-to-strong research.Updated
- WikiSparse autoencoders (interpretability)Sparse autoencoders split an AI model's internal activity into readable features. Their rise, scale-up and limits.Updated
- WikiUK AI Security Institute (AISI)The UK AI Security Institute tests frontier AI models for national-security risks. Its history, tools and key findings.Updated
- AnalysisCan AI safety keep pace with AI capabilities?The central debate in AI safety in 2026: are evaluations, interpretability and oversight keeping up with fast-rising capabilities?Updated
- ExplainerHow AI safety testing works: evals, red teams and thresholdsHow frontier AI models are tested before release: dangerous-capability evals, jailbreak red-teaming, and why testing got harder.Updated
- ExplainerHow mechanistic interpretability works: looking inside AIHow researchers reverse-engineer AI models: features, sparse autoencoders, circuit tracing and the 2026 Jacobian lens.Updated
- ExplainerWhat is AI alignment? Making AI do what we intendAI alignment explained: how labs train models to follow human intent, why it can fail, and what 2024–2026 studies found.Updated