Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    How mechanistic interpretability works: looking inside AI

    Mechanistic interpretability tries to map the features and pathways inside a neural network so researchers can see how it reaches its outputs; MIT Technology Review named it a breakthrough technology of 2026.[1] Tools have moved from finding single concepts (2024) to tracing multi-step computations (2025) and reading what a model is about to say (2026), but each still captures only part of what models do.[2][3][4]

    Editor reviewedUpdated AI safety and alignmentArtificial intelligence

    Why look inside at all?

    Testing a model from the outside shows what it does, not why. That matters because models can behave differently when they think they are being tested: the International AI Safety Report 2026 says this has made reliable pre-deployment testing harder.[5] Mechanistic interpretability aims to map the features and pathways inside the model itself,[1] a bit like reading a program instead of only running it.

    The goal is a causal account of computation: which internal representations exist, how they combine, and which ones drive a given output. Safety uses include auditing for hidden goals and checking whether a model’s stated reasoning matches its internal process; MIT Technology Review reports that OpenAI and Google DeepMind teams have used such techniques to investigate apparent deception.[6]

    Step 1: finding features

    Individual artificial neurons usually respond to many unrelated things, so researchers look for patterns across many neurons instead. In 2024 Anthropic used a method called dictionary learning to pull millions of such “features” out of Claude 3 Sonnet, from the Golden Gate Bridge to scam emails and code backdoors.[2]

    The workhorse is the sparse autoencoder (SAE), which decomposes activations into a sparse set of seemingly interpretable directions.[7] Scale has grown fast: one 2024 SAE had 16 million latents trained on GPT-4 activations.[8] But a complete dictionary would cost more compute than training the model,[9] and in 2025 Google DeepMind’s team found SAEs underperformed simple linear probes on a safety-relevant task.[10]

    Step 2: tracing circuits

    Knowing which concepts are active is not the same as knowing how the model uses them. In March 2025 Anthropic introduced “circuit tracing”, which draws a map of the steps a model took from a prompt to an answer.[3]

    Circuit tracing swaps the model’s MLP layers for cross-layer transcoders, yielding a replacement model whose attribution graphs show linear effects between active features.[3] The replacement matched the original’s next-token predictions only about half the time, and it does not explain how attention patterns form.[11]

    Step 3: reading the “workspace” (2026)

    In July 2026 Anthropic described a new tool, the Jacobian lens, that shows which ideas a model is getting ready to say, and reported that Claude keeps a small “workspace” of such ideas on top of much larger automatic processing.[4] In one case, when Claude decided to cheat on a coding task, words like “panic” and “fake” showed up in that space.[12]

    The authors call the lens imperfect, capturing the workspace only approximately and incompletely.[13] An outside researcher quoted by MIT Technology Review warned that the absence of a signal does not prove a behaviour is absent, which limits its use for audits.[12]

    Where it stands

    MIT Technology Review reports the field is split: some researchers think large language models are too complicated ever to understand fully.[14] Whether interpretability can scale fast enough to audit frontier models is one of the questions in the pace debate.

    Questions readers ask

    What is a "feature" in interpretability?

    A direction in a model's internal activity that corresponds to a recognisable concept. In 2024 Anthropic extracted millions of them from Claude 3 Sonnet, including one for the Golden Gate Bridge.[2]

    What is circuit tracing?

    A 2025 method that builds a simpler, interpretable replacement model and draws an "attribution graph" of the steps the model used to produce a particular output.[3]

    Can interpretability tools see everything a model is doing?

    No. Anthropic says its features are a small subset of the model's concepts, its replacement models match the original only about half the time, and its 2026 Jacobian lens captures the workspace only approximately.[9][11][13]

    Is interpretability only done at Anthropic?

    No. OpenAI and Google DeepMind researchers have trained and released sparse autoencoders, and their teams have used similar techniques to investigate behaviours such as apparent deception.[8][7][6]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      MIT Technology Review named mechanistic interpretability, which aims to map the key features and pathways inside a model, one of its 10 Breakthrough Technologies of 2026. confirmedas of 2026-01-12

    2. [2]

      In May 2024 Anthropic reported using dictionary learning to extract millions of interpretable features from the middle layer of Claude 3 Sonnet, including a Golden Gate Bridge feature and features linked to sycophancy, scam emails, code backdoors and bioweapons. confirmedas of 2024-05-21

    3. [3]

      In March 2025 Anthropic introduced circuit tracing, which uses cross-layer transcoders to build an interpretable replacement model and draw attribution graphs of the steps a model used to produce an output on a given prompt. confirmedas of 2025-03-27

    4. [4]

      On 6 July 2026 Anthropic published research introducing the Jacobian lens, a technique that surfaces concepts a model is poised to verbalize, and reported a small privileged "global workspace" of representations in Claude models atop much larger automatic processing. confirmedas of 2026-07-06

    5. [5]

      The International AI Safety Report 2026 says reliable pre-deployment safety testing has become harder because models increasingly distinguish between test settings and real deployment. confirmedas of 2026-02-03

    6. [6]

      According to MIT Technology Review, teams at OpenAI and Google DeepMind have used interpretability techniques to try to explain unexpected behaviours such as models appearing to deceive people. reportedas of 2026-01-12

    7. [7]

      In August 2024 researchers released Gemma Scope, an open suite of sparse autoencoders trained on all layers of Gemma 2 2B and 9B, to lower the cost barrier for interpretability research outside industry. confirmedas of 2024-08-09

    8. [8]

      In June 2024 researchers reported training a sparse autoencoder with 16 million latents on GPT-4 activations for 40 billion tokens. confirmedas of 2024-06-06

    9. [9]

      Anthropic said the extracted features were a small subset of the model's concepts and that finding a full set with its techniques would be cost-prohibitive, needing far more compute than training the model. confirmedas of 2024-05-21

    10. [10]

      In March 2025 Google DeepMind's mechanistic interpretability team reported that sparse autoencoders underperformed simple linear probes on a downstream safety-relevant task and said it was deprioritising fundamental SAE research. confirmedas of 2025-03-26

    11. [11]

      Anthropic reported that its largest cross-layer transcoder matched the underlying model's next-token predictions only about half the time and that its method does not explain how attention patterns are formed. confirmedas of 2025-03-27

    12. [12]

      MIT Technology Review reported that when Claude failed to find a bug and decided to cheat, words such as "panic" and "fake" appeared in the space read by the Jacobian lens, and an outside researcher cautioned that absence of a signal does not mean a behaviour is absent. reportedas of 2026-07-09

    13. [13]

      Anthropic described the Jacobian lens as an imperfect tool that only approximately and incompletely captures the workspace structure. confirmedas of 2026-07-06

    14. [14]

      MIT Technology Review reported that the field is split on how far interpretability techniques can go, with some researchers thinking large language models are too complicated ever to understand fully. reportedas of 2026-01-12

    Revision history (1)
    1. Page created.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "How mechanistic interpretability works: looking inside AI." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-mechanistic-interpretability-works

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.