Defenses target how models handle malicious requests
The report covered instruction hierarchies, cautionary warnings, and circuit breakers that attempted to interrupt harmful representations. New red-team benchmarks made defenses easier to compare. It also retained the central limitation: success against a collection of known attacks does not guarantee resistance to a new one.

Researchers identify features inside Claude
Anthropic used sparse autoencoders to separate Claude 3 Sonnet’s internal activations into more interpretable components. Adjusting a feature could influence the model’s output, as the widely discussed Golden Gate Bridge example showed. This offered a way to study parts of a model’s behavior without amounting to a complete explanation of the model.

Public institutions develop their own evaluations
The UK AI Safety Institute released Inspect to support evaluations of knowledge, reasoning, and autonomous capabilities. The report described growing international coordination, while questioning how much evaluation work would depend on voluntary cooperation from model developers. Institutional capacity and reliable model access remained separate issues.
