technology
Sparse autoencoders (interpretability)
Also known as SAE, SAEs, dictionary learning, sparse dictionary learning
Sparse autoencoders (SAEs) are an unsupervised method for decomposing a neural network's internal activations into a sparse set of seemingly interpretable features.[1] They powered the 2024 wave of interpretability results, from millions of features in Claude 3 Sonnet to a 16-million-latent SAE trained on GPT-4,[2][3] but in 2025 Google DeepMind's team reported disappointing results on a safety task and shifted effort elsewhere.[4]
Key facts
What they are
Inside a language model, concepts are spread across many artificial neurons at once. A sparse autoencoder is a small auxiliary network trained to rewrite a model’s activations as a combination of a few active “features” drawn from a much larger dictionary. The Gemma Scope authors describe SAEs as an unsupervised method for a sparse decomposition of a network’s latent representations into seemingly interpretable features.[1]
Why they were needed
Individual neurons are hard to interpret because one neuron often responds to several unrelated things. Anthropic‘s interpretability researchers explain this with superposition: a network can represent more features than it has neurons, and this tends to happen when the features useful to a model are sparse in the training data.[5] The authors of “Towards Monosemanticity” chose a sparse autoencoder, which they describe as a weak dictionary learning algorithm, to produce features that are a more monosemantic (single-meaning) unit of analysis than neurons.[6]
The 2023 proof of concept
Published in October 2023, that study took a one-layer transformer with a 512-neuron MLP layer and trained sparse autoencoders on its activations from 8 billion data points, producing dictionaries ranging from 512 to 131,072 features.[7] The result showed that the method could pull relatively interpretable features out of a real, if tiny, model.[7]
The 2024 scale-up
In May 2024 Anthropic reported extracting millions of features from the middle layer of Claude 3 Sonnet. They included a Golden Gate Bridge feature and features tied to safety-relevant behaviour such as sycophantic praise, scam emails, code backdoors and bioweapons.[2] A month later, researchers reported training an SAE with 16 million latents on OpenAI‘s GPT-4 activations over 40 billion tokens.[3] In August 2024 the Gemma Scope release offered open SAEs for every layer of Gemma 2 2B and 9B, aiming to lower the high cost that had kept such work mostly inside industry.[1]
Limits and the 2025 reassessment
Anthropic noted early that its features were only a small subset of the model’s concepts, and that finding a full set with its techniques would be cost-prohibitive, requiring far more compute than training the model.[8] In March 2025 Google DeepMind‘s mechanistic interpretability team reported that SAE-based probes did worse than dense linear probes on a safety-relevant downstream task, and said it was deprioritising fundamental SAE research.[4]
What came next
The field did not abandon sparse dictionaries. Anthropic’s 2025 circuit-tracing work used a related tool, cross-layer transcoders, to build interpretable replacement models and trace computations step by step,[9] and its 2026 Jacobian lens took a different route, reading what a model is poised to say.[10] Anthropic itself calls the Jacobian lens imperfect, capturing the structure it looks for only approximately and incompletely.[11] The UK AI Security Institute counts a model’s internal activations as one of four surfaces that oversight relies on, and warned in 2026 that current oversight methods are likely to erode.[12][13] For the full picture, see how mechanistic interpretability works.
Questions readers ask
What does a sparse autoencoder find?
Features: patterns of activity that correspond to concepts. Anthropic's 2024 work found features for the Golden Gate Bridge, scam emails, sycophancy and code backdoors in Claude 3 Sonnet.[2]
Why are sparse autoencoders controversial?
Finding every feature would cost more than training the model, and in 2025 Google DeepMind's team found SAEs did worse than simple linear probes on a safety-relevant task and deprioritised the research.[8][4]
Can outside researchers use sparse autoencoders?
Yes. Gemma Scope released open SAEs for all layers of Gemma 2 2B and 9B to reduce the high cost of training them.[1]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
In August 2024 researchers released Gemma Scope, an open suite of sparse autoencoders trained on all layers of Gemma 2 2B and 9B, to lower the cost barrier for interpretability research outside industry. confirmedas of 2024-08-09
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 · arXiv · 2024-08-09 (retrieved 2026-10-10)
- [2]
In May 2024 Anthropic reported using dictionary learning to extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including a Golden Gate Bridge feature and features linked to sycophancy, scam emails, code backdoors and bioweapons. confirmedas of 2024-05-21
- Mapping the Mind of a Large Language Model · Anthropic · 2024-05-21 (retrieved 2026-10-10)
- [3]
In June 2024 researchers reported training a sparse autoencoder with 16 million latents on GPT-4 activations for 40 billion tokens. confirmedas of 2024-06-06
- Scaling and evaluating sparse autoencoders · arXiv · 2024-06-06 (retrieved 2026-10-10)
- [4]
In March 2025 Google DeepMind's mechanistic interpretability team reported that sparse autoencoders underperformed simple linear probes on a downstream safety-relevant task and said it was deprioritising fundamental SAE research. confirmedas of 2025-03-26
- Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update · AI Alignment Forum (Google DeepMind mechanistic interpretability team) · 2025-03-26 (retrieved 2026-10-10)
- [5]
Anthropic's interpretability work argues that models represent more features than they have neurons, a phenomenon called superposition, which can arise when useful features are sparse in the training data. confirmedas of 2023-10-04
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning · Anthropic (Transformer Circuits Thread) · 2023-10-04 (retrieved 2026-10-10)
- [6]
The Towards Monosemanticity authors describe a sparse autoencoder as a weak dictionary learning algorithm that yields features that are a more monosemantic unit of analysis than individual neurons. confirmedas of 2023-10-04
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning · Anthropic (Transformer Circuits Thread) · 2023-10-04 (retrieved 2026-10-10)
- [7]
In October 2023 Anthropic's "Towards Monosemanticity" used sparse autoencoders to decompose the 512-neuron MLP layer of a one-layer transformer into up to 131,072 more interpretable features. confirmedas of 2023-10-04
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learning · Anthropic (Transformer Circuits Thread) · 2023-10-04 (retrieved 2026-10-10)
- [8]
Anthropic said the extracted features were a small subset of the model's concepts and that finding a full set with its techniques would be cost-prohibitive, needing far more compute than training the model. confirmedas of 2024-05-21
- Mapping the Mind of a Large Language Model · Anthropic · 2024-05-21 (retrieved 2026-10-10)
- [9]
In March 2025 Anthropic introduced circuit tracing, which uses cross-layer transcoders to build an interpretable replacement model and draw attribution graphs of the steps a model used to produce an output on a given prompt. confirmedas of 2025-03-27
- Circuit Tracing: Revealing Computational Graphs in Language Models · Anthropic (Transformer Circuits Thread) · 2025-03-27 (retrieved 2026-10-10)
- [10]
On 6 July 2026 Anthropic published research introducing the Jacobian lens, a technique that surfaces concepts a model is poised to verbalize, and reported a small privileged "global workspace" of representations in Claude models atop much larger automatic processing. confirmedas of 2026-07-06
- Verbalizable Representations Form a Global Workspace in Language Models · Anthropic (Transformer Circuits Thread) · 2026-07-06 · Introduction; the paper also describes the workspace as "limited in capacity" (retrieved 2026-10-10)
- [11]
Anthropic described the Jacobian lens as an imperfect tool that only approximately and incompletely captures the workspace structure. confirmedas of 2026-07-06
- Verbalizable Representations Form a Global Workspace in Language Models · Anthropic (Transformer Circuits Thread) · 2026-07-06 (retrieved 2026-10-10)
- [12]
The AISI report identifies four oversight surfaces, starting with a model's internal activations, its chain of thought and its external actions. confirmedas of 2026-05-21
- Will it become harder to oversee AI systems? · UK AI Security Institute · 2026-05-21 (retrieved 2026-10-10)
- [13]
AISI's May 2026 report on AI oversight, drawing on 25 expert interviews, concluded that current oversight rests on foundations likely to erode and that emerging methods are not yet mature enough to compensate. confirmedas of 2026-05-21
- Will it become harder to oversee AI systems? · UK AI Security Institute · 2026-05-21 (retrieved 2026-10-10)
- [14]
MIT Technology Review named mechanistic interpretability, which aims to map the key features and pathways inside a model, one of its 10 Breakthrough Technologies of 2026. confirmedas of 2026-01-12
- Mechanistic interpretability: 10 Breakthrough Technologies 2026 · MIT Technology Review · 2026-01-12 (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Added the superposition problem SAEs address and Anthropic's 2023 Towards Monosemanticity results.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Sparse autoencoders (interpretability)." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/sparse-autoencoders
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow mechanistic interpretability works: looking inside AIHow researchers reverse-engineer AI models: features, sparse autoencoders, circuit tracing and the 2026 Jacobian lens.
- ExplainerAI safety and alignment in 2026: a crash courseA sourced crash course on AI safety: alignment, interpretability, evaluations, oversight, safety institutes and the 2026 frontier.
- AnalysisCan AI safety keep pace with AI capabilities?The central debate in AI safety in 2026: are evaluations, interpretability and oversight keeping up with fast-rising capabilities?
- WikiAI controlAI control designs safeguards that hold even if an AI model tries to subvert them: monitoring, trusted editing and sandboxing.
- WikiCAISI / CAISSI: the US government's AI testing centre at NISTThe US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).
- WikiConstitutional AIConstitutional AI trains models with written principles and AI feedback instead of only human labels. How it works and what changed in 2026.