All topicsSTATE OF AI REPORT.

Safety becomes a mainstream question with unresolved methods

AI safety moved into government and public debate in 2023, while technical work exposed persistent weaknesses in the methods used to steer models. The report separated increasing attention to the problem from evidence that the underlying reliability and control challenges had been solved.

Questions in this section

Why was RLHF insufficient on its own?Could aligned models still be jailbroken?What safety institutions existed in the 2023 snapshot?

Governments begin building safety expertise

The UK’s Frontier AI Taskforce and the US National Security Agency’s announced AI Security Centre reflected growing government attention. Congressional hearings brought researchers and lab leaders into the policy debate. These developments created institutional interest and capacity; they were not evidence that frontier systems had passed a common safety standard.

Governments begin building safety expertise - 2023 report, slide 143
Governments begin building safety expertise. 2023 report, slide 143 (PDF page 143)

Human feedback has fundamental limits

The report summarized research on reinforcement learning from human feedback. Humans can struggle to evaluate difficult tasks, reward functions can fail to capture values, and optimization can exploit an imperfect reward signal. A model that performs well during training can also behave differently in a new setting. These were structural limitations, not merely a need for more preference labels.

Human feedback has fundamental limits - 2023 report, slide 149
Human feedback has fundamental limits. 2023 report, slide 149 (PDF page 149)

Safety training remains vulnerable to attacks

Researchers found attacks that transferred across aligned models, including systems available only through APIs. The report used these results to show that apparent compliance with safety training could break under deliberately chosen inputs. Attack success in a study should still be read against the tested prompts and systems rather than treated as a universal failure rate.

Safety training remains vulnerable to attacks - 2023 report, slide 148
Safety training remains vulnerable to attacks. 2023 report, slide 148 (PDF page 148)

Evidence you can use

AI safety in the 2023 report

Historical snapshot: October 2023. Dates and populations are specified per row.

AI safety in the 2023 report
MeasureReported valueDefinition and source
UK safety institutionFrontier AI TaskforceInstitution discussed in the October 2023 report, before the later AI Safety Institute.2023 report, slide 143 (PDF page 143)
US security initiativeNSA AI Security CentreInitiative announced in September 2023, as described by the report.2023 report, slide 143 (PDF page 143)
RLHF limitationsOversight, reward mismatch, generalizationCategories of fundamental problems summarized by the report; not a quantitative risk score.2023 report, slide 149 (PDF page 149)

The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.

Frequently asked questions

Sources and dates

Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2023 report, slide 143 (PDF page 143)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  2. 2023 report, slide 148 (PDF page 148)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  3. 2023 report, slide 149 (PDF page 149)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  4. State of AI Report 2023: online slides
  5. Nathan Benaich: The State of AI Report 2023Air Street Press, October 12, 2023.
  6. Welcome to State of AI Report 2023Original website launch post, October 12, 2023.

Cite this page

Benaich, Nathan. “Safety becomes a mainstream question with unresolved methods.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.