
How to read this finding
The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.
Human feedback can be unreliable on difficult tasks, reward functions imperfectly represent the intended goals, and learned behavior can fail to generalize. The report treated these as fundamental limitations of relying on RLHF alone.

The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.
Evidence you can use
Historical snapshot: October 2023. Dates and populations are specified per row.
| Measure | Reported value | Definition and source |
|---|---|---|
| UK safety institution | Frontier AI Taskforce | Institution discussed in the October 2023 report, before the later AI Safety Institute.2023 report, slide 143 (PDF page 143) |
| US security initiative | NSA AI Security Centre | Initiative announced in September 2023, as described by the report.2023 report, slide 143 (PDF page 143) |
| RLHF limitations | Oversight, reward mismatch, generalization | Categories of fundamental problems summarized by the report; not a quantitative risk score.2023 report, slide 149 (PDF page 149) |
The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.
Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan. “Why was RLHF insufficient on its own?.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.