Better language-model results increasingly depended on larger models, more data, and more compute. The 2020 report asked who could afford that trajectory. It also showed that rising frontier budgets could coexist with falling compute needs for a fixed level of performance.
The report cited an expert estimate of roughly $10 million to train GPT-3. It was an estimate, not an audited bill from OpenAI. The broader point was that the cost of leading experiments could narrow the set of organizations able to attempt them.
A fixed task can get cheaper as the frontier gets larger
For a fixed level of ImageNet performance, the report cited a halving of required training compute every 16 months since 2012. This measured algorithmic efficiency on one task. It did not contradict rising spending on much larger frontier models.
The report’s efficiency result asked how much compute was needed to reach a fixed ImageNet performance level. Frontier-model spending asked a different question: how far could performance be pushed? Falling cost for an existing target could therefore coexist with rising expenditure on more ambitious models.
Evidence you can use
Compute in the 2020 report
Historical snapshot: October 2020. Dates and populations are specified per row.
The two figures describe different quantities: a model-specific cost estimate and a historical fixed-performance efficiency trend. Neither should be extrapolated as a universal rule or used as a current price.
For fixed ImageNet performance, required compute was falling. At the frontier, organizations were also choosing to train larger models, so total spending could rise at the same time.
No. It was an expert estimate cited by the report. The page discussed cost assumptions and comparisons rather than presenting a verified accounting record for the training run.
The cited number concerned training. It should not be treated as a complete total for research salaries, failed experiments, deployment, or subsequent service operation.
Larger models increased the scale of the training problem. The report used model size to motivate the cost challenge, while its estimates also depended on assumptions about hardware and training.
The comparison held a performance level fixed and tracked the compute needed to reach it. It therefore measured efficiency improvements rather than the maximum achievable accuracy each year.
Within the cited ImageNet comparison, the compute required for the same performance level fell by half about every sixteen months. The result was specific to that task and analysis.
Not necessarily. Doing an existing task more efficiently and spending more compute on larger ambitions are compatible. The fixed-performance comparison did not measure total industry demand.
Historical snapshot published October 1, 2020. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
2020 report, PDF page 17Original 2020 report. This edition has no printed slide numbers. References use one-based PDF pages.
2020 report, PDF page 22Original 2020 report. This edition has no printed slide numbers. References use one-based PDF pages.
Benaich, Nathan, and Ian Hogarth. “Scaling creates a growing resource divide.” State of AI Report 2020. Historical report snapshot; web edition prepared 2026-10-11.