Governments begin building safety expertise
The UK’s Frontier AI Taskforce and the US National Security Agency’s announced AI Security Centre reflected growing government attention. Congressional hearings brought researchers and lab leaders into the policy debate. These developments created institutional interest and capacity; they were not evidence that frontier systems had passed a common safety standard.

Human feedback has fundamental limits
The report summarized research on reinforcement learning from human feedback. Humans can struggle to evaluate difficult tasks, reward functions can fail to capture values, and optimization can exploit an imperfect reward signal. A model that performs well during training can also behave differently in a new setting. These were structural limitations, not merely a need for more preference labels.

Safety training remains vulnerable to attacks
Researchers found attacks that transferred across aligned models, including systems available only through APIs. The report used these results to show that apparent compliance with safety training could break under deliberately chosen inputs. Attack success in a study should still be read against the tested prompts and systems rather than treated as a universal failure rate.
