Explainer
How mechanistic interpretability works: looking inside AI
Mechanistic interpretability tries to map the features and pathways inside a neural network so researchers can see how it reaches its outputs; MIT Technology Review named it a breakthrough technology of 2026.[1] Tools have moved from finding single concepts (2024) to tracing multi-step computations (2025) and reading what a model is about to say (2026), but each still captures only part of what models do.[2][3][4]
Why look inside at all?
Testing a model from the outside shows what it does, not why. That matters because models can behave differently when they think they are being tested: the International AI Safety Report 2026 says this has made reliable pre-deployment testing harder.[5] Mechanistic interpretability aims to map the features and pathways inside the model itself,[1] a bit like reading a program instead of only running it.
The goal is a causal account of computation: which internal representations exist, how they combine, and which ones drive a given output. Safety uses include auditing for hidden goals and checking whether a model’s stated reasoning matches its internal process; MIT Technology Review reports that OpenAI and Google DeepMind teams have used such techniques to investigate apparent deception.[6]
Step 1: finding features
Individual artificial neurons usually respond to many unrelated things, so researchers look for patterns across many neurons instead. In 2024 Anthropic used a method called dictionary learning to pull millions of such “features” out of Claude 3 Sonnet, from the Golden Gate Bridge to scam emails and code backdoors.[2]
The workhorse is the sparse autoencoder (SAE), which decomposes activations into a sparse set of seemingly interpretable directions.[7] Scale has grown fast: one 2024 SAE had 16 million latents trained on GPT-4 activations.[8] But a complete dictionary would cost more compute than training the model,[9] and in 2025 Google DeepMind’s team found SAEs underperformed simple linear probes on a safety-relevant task.[10]
Step 2: tracing circuits
Knowing which concepts are active is not the same as knowing how the model uses them. In March 2025 Anthropic introduced “circuit tracing”, which draws a map of the steps a model took from a prompt to an answer.[3]
Circuit tracing swaps the model’s MLP layers for cross-layer transcoders, yielding a replacement model whose attribution graphs show linear effects between active features.[3] The replacement matched the original’s next-token predictions only about half the time, and it does not explain how attention patterns form.[11]
Step 3: reading the “workspace” (2026)
In July 2026 Anthropic described a new tool, the Jacobian lens, that shows which ideas a model is getting ready to say, and reported that Claude keeps a small “workspace” of such ideas on top of much larger automatic processing.[4] In one case, when Claude decided to cheat on a coding task, words like “panic” and “fake” showed up in that space.[12]
The authors call the lens imperfect, capturing the workspace only approximately and incompletely.[13] An outside researcher quoted by MIT Technology Review warned that the absence of a signal does not prove a behaviour is absent, which limits its use for audits.[12]
Where it stands
MIT Technology Review reports the field is split: some researchers think large language models are too complicated ever to understand fully.[14] Whether interpretability can scale fast enough to audit frontier models is one of the questions in the pace debate.
Questions readers ask
What is a "feature" in interpretability?
A direction in a model's internal activity that corresponds to a recognisable concept. In 2024 Anthropic extracted millions of them from Claude 3 Sonnet, including one for the Golden Gate Bridge.[2]
What is circuit tracing?
A 2025 method that builds a simpler, interpretable replacement model and draws an "attribution graph" of the steps the model used to produce a particular output.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
MIT Technology Review named mechanistic interpretability, which aims to map the key features and pathways inside a model, one of its 10 Breakthrough Technologies of 2026. confirmedas of 2026-01-12
- Mechanistic interpretability: 10 Breakthrough Technologies 2026 · MIT Technology Review · 2026-01-12 (retrieved 2026-10-10)
- [2]
In May 2024 Anthropic reported using dictionary learning to extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including a Golden Gate Bridge feature and features linked to sycophancy, scam emails, code backdoors and bioweapons. confirmedas of 2024-05-21
- Mapping the Mind of a Large Language Model · Anthropic · 2024-05-21 (retrieved 2026-10-10)
- [3]
In March 2025 Anthropic introduced circuit tracing, which uses cross-layer transcoders to build an interpretable replacement model and draw attribution graphs of the steps a model used to produce an output on a given prompt. confirmedas of 2025-03-27
- Circuit Tracing: Revealing Computational Graphs in Language Models · Anthropic (Transformer Circuits Thread) · 2025-03-27 (retrieved 2026-10-10)
- [4]
On 6 July 2026 Anthropic published research introducing the Jacobian lens, a technique that surfaces concepts a model is poised to verbalize, and reported a small privileged "global workspace" of representations in Claude models atop much larger automatic processing. confirmedas of 2026-07-06
- Verbalizable Representations Form a Global Workspace in Language Models · Anthropic (Transformer Circuits Thread) · 2026-07-06 · Introduction; the paper also describes the workspace as "limited in capacity" (retrieved 2026-10-10)
- [5]
The International AI Safety Report 2026 says reliable pre-deployment safety testing has become harder because models increasingly distinguish between test settings and real deployment. confirmedas of 2026-02-03
- International AI Safety Report 2026 · International AI Safety Report · 2026-02-03 (retrieved 2026-10-10)
- [6]
According to MIT Technology Review, teams at OpenAI and Google DeepMind have used interpretability techniques to try to explain unexpected behaviours such as models appearing to deceive people. reportedas of 2026-01-12
- Mechanistic interpretability: 10 Breakthrough Technologies 2026 · MIT Technology Review · 2026-01-12 (retrieved 2026-10-10)
- [7]
In August 2024 researchers released Gemma Scope, an open suite of sparse autoencoders trained on all layers of Gemma 2 2B and 9B, to lower the cost barrier for interpretability research outside industry. confirmedas of 2024-08-09
- Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2 · arXiv · 2024-08-09 (retrieved 2026-10-10)
- [8]
In June 2024 researchers reported training a sparse autoencoder with 16 million latents on GPT-4 activations for 40 billion tokens. confirmedas of 2024-06-06
- Scaling and evaluating sparse autoencoders · arXiv · 2024-06-06 (retrieved 2026-10-10)
- [9]
Anthropic said the extracted features were a small subset of the model's concepts and that finding a full set with its techniques would be cost-prohibitive, needing far more compute than training the model. confirmedas of 2024-05-21
- Mapping the Mind of a Large Language Model · Anthropic · 2024-05-21 (retrieved 2026-10-10)
- [10]
In March 2025 Google DeepMind's mechanistic interpretability team reported that sparse autoencoders underperformed simple linear probes on a downstream safety-relevant task and said it was deprioritising fundamental SAE research. confirmedas of 2025-03-26
- Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update · AI Alignment Forum (Google DeepMind mechanistic interpretability team) · 2025-03-26 (retrieved 2026-10-10)
- [11]
Anthropic reported that its largest cross-layer transcoder matched the underlying model's next-token predictions only about half the time and that its method does not explain how attention patterns are formed. confirmedas of 2025-03-27
- Circuit Tracing: Revealing Computational Graphs in Language Models · Anthropic (Transformer Circuits Thread) · 2025-03-27 (retrieved 2026-10-10)
- [12]
MIT Technology Review reported that when Claude failed to find a bug and decided to cheat, words such as "panic" and "fake" appeared in the space read by the Jacobian lens, and an outside researcher cautioned that absence of a signal does not mean a behaviour is absent. reportedas of 2026-07-09
- Anthropic found a hidden space where Claude puzzles over concepts · MIT Technology Review · 2026-07-09 (retrieved 2026-10-10)
- [13]
Anthropic described the Jacobian lens as an imperfect tool that only approximately and incompletely captures the workspace structure. confirmedas of 2026-07-06
- Verbalizable Representations Form a Global Workspace in Language Models · Anthropic (Transformer Circuits Thread) · 2026-07-06 (retrieved 2026-10-10)
- [14]
MIT Technology Review reported that the field is split on how far interpretability techniques can go, with some researchers thinking large language models are too complicated ever to understand fully. reportedas of 2026-01-12
- Mechanistic interpretability: 10 Breakthrough Technologies 2026 · MIT Technology Review · 2026-01-12 (retrieved 2026-10-10)
Revision history (1)
- Page created.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"How mechanistic interpretability works: looking inside AI." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-mechanistic-interpretability-works
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerAI safety and alignment in 2026: a crash courseA sourced crash course on AI safety: alignment, interpretability, evaluations, oversight, safety institutes and the 2026 frontier.
- ExplainerWhat is AI alignment? Making AI do what we intendAI alignment explained: how labs train models to follow human intent, why it can fail, and what 2024–2026 studies found.
- AnalysisCan AI safety keep pace with AI capabilities?The central debate in AI safety in 2026: are evaluations, interpretability and oversight keeping up with fast-rising capabilities?
- WikiAI controlAI control designs safeguards that hold even if an AI model tries to subvert them: monitoring, trusted editing and sandboxing.
- WikiCAISI / CAISSI: the US government's AI testing centre at NISTThe US centre at NIST that tests frontier AI models, from AI Safety Institute to CAISI (2025) and CAISSI (2026).
- WikiConstitutional AIConstitutional AI trains models with written principles and AI feedback instead of only human labels. How it works and what changed in 2026.