All topicsSTATE OF AI REPORT.

Safety work moves into testing and model internals

The 2024 report documented practical work on evaluating and controlling models alongside a changing public debate about AI risk. New defenses addressed specific failure modes, and interpretability researchers identified features inside a frontier model. Neither established that complex systems had become fully understood or reliably safe.

Questions in this section

What did Anthropic’s interpretability work show in 2024?Were models resistant to jailbreaks?What was Inspect?

Defenses target how models handle malicious requests

The report covered instruction hierarchies, cautionary warnings, and circuit breakers that attempted to interrupt harmful representations. New red-team benchmarks made defenses easier to compare. It also retained the central limitation: success against a collection of known attacks does not guarantee resistance to a new one.

Defenses target how models handle malicious requests - 2024 report, slide 181
Defenses target how models handle malicious requests. 2024 report, slide 181 (PDF page 182)

Researchers identify features inside Claude

Anthropic used sparse autoencoders to separate Claude 3 Sonnet’s internal activations into more interpretable components. Adjusting a feature could influence the model’s output, as the widely discussed Golden Gate Bridge example showed. This offered a way to study parts of a model’s behavior without amounting to a complete explanation of the model.

Researchers identify features inside Claude - 2024 report, slide 197
Researchers identify features inside Claude. 2024 report, slide 197 (PDF page 198)

Public institutions develop their own evaluations

The UK AI Safety Institute released Inspect to support evaluations of knowledge, reasoning, and autonomous capabilities. The report described growing international coordination, while questioning how much evaluation work would depend on voluntary cooperation from model developers. Institutional capacity and reliable model access remained separate issues.

Public institutions develop their own evaluations - 2024 report, slide 178
Public institutions develop their own evaluations. 2024 report, slide 178 (PDF page 179)

Evidence you can use

AI safety in the 2024 report

Historical snapshot: October 2024. Dates and populations are specified per row.

AI safety in the 2024 report
MeasureReported valueDefinition and source
Instruction hierarchyDeployed in GPT-4o MiniDefense described as prioritizing instructions rather than treating all text as equally authoritative.2024 report, slide 181 (PDF page 182)
Interpretability model studiedClaude 3 SonnetModel whose activations were decomposed into interpretable features in the cited Anthropic work.2024 report, slide 197 (PDF page 198)
Public evaluation frameworkInspectFramework released by the UK AI Safety Institute.2024 report, slide 178 (PDF page 179)

These rows identify methods, models, and institutions rather than comparable safety scores. The report did not claim that jailbreak defenses were universal or that interpreting selected features explained all model behavior.

Frequently asked questions

Sources and dates

Historical snapshot published October 10, 2024. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2024 report, slide 178 (PDF page 179)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  2. 2024 report, slide 181 (PDF page 182)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  3. 2024 report, slide 197 (PDF page 198)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  4. State of AI Report 2024: online slides
  5. Nathan Benaich: The State of AI Report 2024Air Street Press, October 10, 2024.
  6. Welcome to State of AI Report 2024Original website launch post, October 10, 2024.

Cite this page

Benaich, Nathan. “Safety work moves into testing and model internals.” State of AI Report 2024. Historical report snapshot; web edition prepared 2026-10-10.