Alignment research is small relative to capability development
By Nathan Benaich and Ian Hogarth · 2021 report
The 2021 report asked whether enough research effort was going toward ensuring that increasingly capable systems behaved as intended. It separated long-term alignment from the broader field of near-term AI safety and documented a limited number of dedicated researchers at selected organizations.
The report’s primary research identified fewer than 100 researchers working on long-term alignment across seven leading organizations. The definition excluded broader work on near-term safety, so the figure should not be read as the total number of people addressing all AI risks.
Capability gains do not guarantee trustworthy answers
The TruthfulQA comparison showed a substantial gap between language models and humans on questions designed to elicit misconceptions. This provided a concrete example of why larger or more fluent systems still needed evaluation beyond conventional performance benchmarks.
A narrow staffing estimate still exposed a priority gap
The alignment count focused on a selected set of organizations and a particular long-term research concern. It could not describe the whole safety field. Even with that limitation, the report used the small scale of the effort to ask how resources were divided between improving capabilities and understanding their consequences.
Evidence you can use
AI safety in the 2021 report
Historical snapshot: October 2021. Dates and populations are specified per row.
The staffing estimate has a narrow organizational and research scope. TruthfulQA measures a particular failure mode; it is not a comprehensive measure of alignment or all model behavior.
It focused on long-term AI alignment at seven selected organizations. It did not count every person working on fairness, robustness, security, or other safety-related topics.
The organization list and research definition were limited. The report’s estimate illustrated the scale of a selected effort rather than exhaustively identifying everyone working on AI safety.
It asked how increasingly powerful systems could be made to work in ways that benefit humanity. The discussion looked beyond present model accuracy toward the consequences of more capable systems.
The report placed a 94% human truthfulness baseline alongside the best-model result of 58%. Both figures referred to the benchmark’s selected questions.
Human text contains popular misconceptions as well as accurate statements. The report argued that models could reproduce those patterns, making natural-sounding answers an unreliable proxy for truth.
No. TruthfulQA tested a present model behavior on a defined question set. The staffing discussion concerned a broader long-term alignment agenda; neither measure fully captured the other.
Historical snapshot published October 12, 2021. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “Alignment research is small relative to capability development.” State of AI Report 2021. Historical report snapshot; web edition prepared 2026-10-11.