<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>ContentLora: AI safety and alignment</title><description>New and updated AI safety and alignment pages on ContentLora.</description><link>https://contentlora.com/</link><language>en</language><atom:link href="https://contentlora.com/topics/ai-safety/feed.xml" rel="self" type="application/rss+xml"/><item><title>AI control</title><link>https://contentlora.com/wiki/ai-control/</link><guid isPermaLink="true">https://contentlora.com/wiki/ai-control/</guid><description>AI control designs safeguards that hold even if an AI model tries to subvert them: monitoring, trusted editing and sandboxing.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category></item><item><title>CAISI / CAISSI: the US government&apos;s AI testing centre at NIST</title><link>https://contentlora.com/wiki/caisi/</link><guid isPermaLink="true">https://contentlora.com/wiki/caisi/</guid><description>The US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>policy</category><category>ai</category></item><item><title>Constitutional AI</title><link>https://contentlora.com/wiki/constitutional-ai/</link><guid isPermaLink="true">https://contentlora.com/wiki/constitutional-ai/</guid><description>Constitutional AI trains models with written principles and AI feedback instead of only human labels. How it works and what changed in 2026.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category></item><item><title>Frontier safety frameworks (responsible scaling policies)</title><link>https://contentlora.com/wiki/frontier-safety-frameworks/</link><guid isPermaLink="true">https://contentlora.com/wiki/frontier-safety-frameworks/</guid><description>Frontier safety frameworks are AI companies&apos; if-then rules for dangerous capabilities. How they work, who has one, and 2026 changes.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category><category>policy</category></item><item><title>Reinforcement learning from human feedback (RLHF)</title><link>https://contentlora.com/wiki/rlhf/</link><guid isPermaLink="true">https://contentlora.com/wiki/rlhf/</guid><description>RLHF trains AI models on human preference rankings. Where it came from, why it became the standard, and its limits.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category></item><item><title>Scalable oversight</title><link>https://contentlora.com/wiki/scalable-oversight/</link><guid isPermaLink="true">https://contentlora.com/wiki/scalable-oversight/</guid><description>Scalable oversight asks how humans can supervise AI that outperforms them. Debate, AI-assisted judging and weak-to-strong research.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category></item><item><title>Sparse autoencoders (interpretability)</title><link>https://contentlora.com/wiki/sparse-autoencoders/</link><guid isPermaLink="true">https://contentlora.com/wiki/sparse-autoencoders/</guid><description>Sparse autoencoders split an AI model&apos;s internal activity into readable features. Their rise, scale-up and limits.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category></item><item><title>UK AI Security Institute (AISI)</title><link>https://contentlora.com/wiki/uk-ai-security-institute/</link><guid isPermaLink="true">https://contentlora.com/wiki/uk-ai-security-institute/</guid><description>The UK AI Security Institute tests frontier AI models for national-security risks. Its history, tools and key findings.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Wiki</category><category>ai-safety</category><category>ai</category><category>policy</category></item><item><title>AI safety tracker: alignment, interpretability and evals in 2026</title><link>https://contentlora.com/events/ai-safety-tracker/</link><guid isPermaLink="true">https://contentlora.com/events/ai-safety-tracker/</guid><description>Live tracker of AI safety milestones: interpretability results, evaluations, safety frameworks, incidents and institutes.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Developing</category><category>ai-safety</category><category>ai</category><category>policy</category></item><item><title>Can AI safety keep pace with AI capabilities?</title><link>https://contentlora.com/analysis/ai-safety-keeping-pace-debate/</link><guid isPermaLink="true">https://contentlora.com/analysis/ai-safety-keeping-pace-debate/</guid><description>The central debate in AI safety in 2026: are evaluations, interpretability and oversight keeping up with fast-rising capabilities?</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Analysis</category><category>ai-safety</category><category>ai</category></item><item><title>AI safety and alignment in 2026: a crash course</title><link>https://contentlora.com/explain/ai-safety/</link><guid isPermaLink="true">https://contentlora.com/explain/ai-safety/</guid><description>A sourced crash course on AI safety: alignment, interpretability, evaluations, oversight, safety institutes and the 2026 frontier.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Explainer</category><category>ai-safety</category><category>ai</category><category>policy</category></item><item><title>How AI safety testing works: evals, red teams and thresholds</title><link>https://contentlora.com/explain/how-ai-safety-testing-works/</link><guid isPermaLink="true">https://contentlora.com/explain/how-ai-safety-testing-works/</guid><description>How frontier AI models are tested before release: dangerous-capability evals, jailbreak red-teaming, and why testing got harder.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Explainer</category><category>ai-safety</category><category>ai</category></item><item><title>How mechanistic interpretability works: looking inside AI</title><link>https://contentlora.com/explain/how-mechanistic-interpretability-works/</link><guid isPermaLink="true">https://contentlora.com/explain/how-mechanistic-interpretability-works/</guid><description>How researchers reverse-engineer AI models: features, sparse autoencoders, circuit tracing and the 2026 Jacobian lens.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Explainer</category><category>ai-safety</category><category>ai</category></item><item><title>What is AI alignment? Making AI do what we intend</title><link>https://contentlora.com/explain/what-is-ai-alignment/</link><guid isPermaLink="true">https://contentlora.com/explain/what-is-ai-alignment/</guid><description>AI alignment explained: how labs train models to follow human intent, why it can fail, and what 2024–2026 studies found.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><category>Explainer</category><category>ai-safety</category><category>ai</category></item></channel></rss>