Safety gains its own section and a larger research community
By Nathan Benaich and Ian Hogarth · 2022 report
The 2022 report made safety a standalone section for the first time. It documented growing interest in the risks of highly capable systems and explained methods for improving language-model behavior. More research activity was a development to track, not evidence that the underlying problems were solved.
The report estimated around 300 researchers working full-time on AI safety and described new organizations and training programs. This remained small relative to capability research. Its scope also differed from the previous year’s narrow count of alignment researchers at seven organizations.
Human feedback shapes instruction-following behavior
Reinforcement learning from human feedback uses people’s rankings of model outputs to train a preference model, then uses that signal to fine-tune the language model. The report described improved instruction-following from InstructGPT and related work. Better preferences on tested outputs were not a guarantee of safe behavior in every setting.
The report described an NYU approach that used human feedback written in natural language to improve a model directly. With 100 feedback samples, the researchers improved GPT-3 on a summarization task to the reported human level. The example raised a practical question alongside RLHF: how much information was being discarded when detailed human criticism was reduced to a preference or score?
The researcher estimate is not directly comparable with the narrower 2021 alignment count. Fine-tuning compute excludes the cost of pretraining, and improved instruction following does not establish comprehensive alignment.
It used human rankings to learn a preference signal and fine-tune a language model toward more useful behavior. The report described practical improvements, not a complete solution to alignment.
Not cleanly. The 2021 figure focused on long-term alignment at seven selected organizations; the 2022 estimate used a broader AI-safety framing. Different scopes prevent a simple like-for-like growth calculation.
No. The report described greater attention and more researchers while still characterizing the field as neglected relative to the scale of capability development and its risks.
It provided judgments about model responses that could guide fine-tuning toward human preferences. The approach aimed to improve how an existing language model behaved when following instructions.
It compared the fine-tuning compute described for InstructGPT with GPT-3 pretraining compute. It did not measure the complete cost of developing, evaluating, or operating the model.
It quantified the human-feedback effort cited in the InstructGPT discussion. That labor input was separate from the compute used to train or fine-tune the model.
The report described an NYU method that fine-tuned GPT-3 using feedback written in natural language. With 100 feedback samples, it reached reported human-level performance on a summarization task. The result concerned that task and evaluation, rather than a general guarantee of aligned behavior.
Historical snapshot published October 11, 2022. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “Safety gains its own section and a larger research community.” State of AI Report 2022. Historical report snapshot; web edition prepared 2026-10-11.