
How to read this finding
These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.
Sparse autoencoders identified interpretable components in Claude 3 Sonnet’s activations, and adjusting features could influence outputs. This was progress in studying model internals, not a complete account of how the model worked.

These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.
Evidence you can use
Historical snapshot: October 2024. Dates and populations are specified per row.
| Measure | Reported value | Definition and source |
|---|---|---|
| Instruction hierarchy | Deployed in GPT-4o Mini | Defense described as prioritizing instructions rather than treating all text as equally authoritative.2024 report, slide 181 (PDF page 182) |
| Interpretability model studied | Claude 3 Sonnet | Model whose activations were decomposed into interpretable features in the cited Anthropic work.2024 report, slide 197 (PDF page 198) |
| Public evaluation framework | Inspect | Framework released by the UK AI Safety Institute.2024 report, slide 178 (PDF page 179) |
These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.
Historical snapshot published October 10, 2024. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan. “What did Anthropic’s interpretability work show in 2024?.” State of AI Report 2024. Historical report snapshot; web edition prepared 2026-10-10.