All topicsSTATE OF AI REPORT.

Why was RLHF insufficient on its own?

Human feedback can be unreliable on difficult tasks, reward functions imperfectly represent the intended goals, and learned behavior can fail to generalize. The report treated these as fundamental limitations of relying on RLHF alone.

Why was RLHF insufficient on its own? - 2023 report, slide 149
Why was RLHF insufficient on its own?. 2023 report, slide 149 (PDF page 149)

How to read this finding

The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.

Read the full AI safety section

Evidence you can use

AI safety in the 2023 report

Historical snapshot: October 2023. Dates and populations are specified per row.

AI safety in the 2023 report
MeasureReported valueDefinition and source
UK safety institutionFrontier AI TaskforceInstitution discussed in the October 2023 report, before the later AI Safety Institute.2023 report, slide 143 (PDF page 143)
US security initiativeNSA AI Security CentreInitiative announced in September 2023, as described by the report.2023 report, slide 143 (PDF page 143)
RLHF limitationsOversight, reward mismatch, generalizationCategories of fundamental problems summarized by the report; not a quantitative risk score.2023 report, slide 149 (PDF page 149)

The report combines institutional developments with technical studies. Their presence is not evidence of a shared safety threshold. Attack results depend on the selected models, prompts, and evaluation conditions.

Sources and dates

Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2023 report, slide 143 (PDF page 143)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  2. 2023 report, slide 148 (PDF page 148)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  3. 2023 report, slide 149 (PDF page 149)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  4. State of AI Report 2023: online slides
  5. Nathan Benaich: The State of AI Report 2023Air Street Press, October 12, 2023.
  6. Welcome to State of AI Report 2023Original website launch post, October 12, 2023.

Cite this page

Benaich, Nathan. “Why was RLHF insufficient on its own?.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.