All topicsSTATE OF AI REPORT. 2022

Safety gains its own section and a larger research community

The 2022 report made safety a standalone section for the first time. It documented growing interest in the risks of highly capable systems and explained methods for improving language-model behavior. More research activity was a development to track, not evidence that the underlying problems were solved.

Questions in this section

How large was AI safety research in the 2022 report?What did RLHF do?Was the approximately 300-person safety estimate directly comparable with 2021?Did growing safety staffing mean the report considered the field adequately resourced?Why was human feedback used in InstructGPT?What did the less-than-2% compute figure compare?What did the 20,000 hours quantify?Could models learn from written feedback instead of only preference scores?

A small field attracts more people

The report estimated around 300 researchers working full-time on AI safety and described new organizations and training programs. This remained small relative to capability research. Its scope also differed from the previous year’s narrow count of alignment researchers at seven organizations.

A small field attracts more people - 2022 report, slide 98
A small field attracts more people. 2022 report, slide 98 (PDF page 98)

Human feedback shapes instruction-following behavior

Reinforcement learning from human feedback uses people’s rankings of model outputs to train a preference model, then uses that signal to fine-tune the language model. The report described improved instruction-following from InstructGPT and related work. Better preferences on tested outputs were not a guarantee of safe behavior in every setting.

Human feedback shapes instruction-following behavior - 2022 report, slide 100
Human feedback shapes instruction-following behavior. 2022 report, slide 100 (PDF page 100)

Written feedback offered a richer training signal

The report described an NYU approach that used human feedback written in natural language to improve a model directly. With 100 feedback samples, the researchers improved GPT-3 on a summarization task to the reported human level. The example raised a practical question alongside RLHF: how much information was being discarded when detailed human criticism was reduced to a preference or score?

Written feedback offered a richer training signal - 2022 report, slide 101
Written feedback offered a richer training signal. 2022 report, slide 101 (PDF page 101)

Evidence you can use

AI safety in the 2022 report

Historical snapshot: October 2022. Dates and populations are specified per row.

AI safety in the 2022 report
MeasureReported valueDefinition and source
Full-time AI safety researchersApproximately 300Estimate reported in 2022; scope differs from the 2021 selected-organization count.2022 report, slide 98 (PDF page 98)
InstructGPT fine-tuning computeLess than 2% of GPT-3 pretrainingRelative compute figure reported for the fine-tuning stage.2022 report, slide 100 (PDF page 100)
Human feedback effort20,000 hoursEffort reported alongside the InstructGPT discussion.2022 report, slide 100 (PDF page 100)

The researcher estimate is not directly comparable with the narrower 2021 alignment count. Fine-tuning compute excludes the cost of pretraining, and improved instruction following does not establish comprehensive alignment.

Frequently asked questions

Sources and dates

Historical snapshot published October 11, 2022. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2022 report, slide 98 (PDF page 98)Original 2022 report. Printed slide numbers match PDF pages.
  2. 2022 report, slide 100 (PDF page 100)Original 2022 report. Printed slide numbers match PDF pages.
  3. 2022 report, slide 101 (PDF page 101)Original 2022 report. Printed slide numbers match PDF pages.
  4. State of AI Report 2022: online slides
  5. Air Street Press launch essayOctober 11, 2022.
  6. Welcome to State of AI Report 2022Original website launch post, October 11, 2022.

Cite this page

Benaich, Nathan, and Ian Hogarth. “Safety gains its own section and a larger research community.” State of AI Report 2022. Historical report snapshot; web edition prepared 2026-10-11.