All topicsSTATE OF AI REPORT.

What did Anthropic’s interpretability work show in 2024?

Sparse autoencoders identified interpretable components in Claude 3 Sonnet’s activations, and adjusting features could influence outputs. This was progress in studying model internals, not a complete account of how the model worked.

What did Anthropic’s interpretability work show in 2024? - 2024 report, slide 197
What did Anthropic’s interpretability work show in 2024?. 2024 report, slide 197 (PDF page 198)

How to read this finding

These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.

Read the full AI safety section

Evidence you can use

AI safety in the 2024 report

Historical snapshot: October 2024. Dates and populations are specified per row.

AI safety in the 2024 report
MeasureReported valueDefinition and source
Instruction hierarchyDeployed in GPT-4o MiniDefense described as prioritizing instructions rather than treating all text as equally authoritative.2024 report, slide 181 (PDF page 182)
Interpretability model studiedClaude 3 SonnetModel whose activations were decomposed into interpretable features in the cited Anthropic work.2024 report, slide 197 (PDF page 198)
Public evaluation frameworkInspectFramework released by the UK AI Safety Institute.2024 report, slide 178 (PDF page 179)

These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.

Sources and dates

Historical snapshot published October 10, 2024. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2024 report, slide 178 (PDF page 179)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  2. 2024 report, slide 181 (PDF page 182)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  3. 2024 report, slide 197 (PDF page 198)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  4. State of AI Report 2024: online slides
  5. Nathan Benaich: The State of AI Report 2024Air Street Press, October 10, 2024.
  6. Welcome to State of AI Report 2024Original website launch post, October 10, 2024.

Cite this page

Benaich, Nathan. “What did Anthropic’s interpretability work show in 2024?.” State of AI Report 2024. Historical report snapshot; web edition prepared 2026-10-10.