The 2020 report described language models that could generate, summarize, translate, and turn text into code. Scaling offered a route to better results across tasks, but the ability to inspect and reproduce research lagged behind the headline capabilities.
GPT-3, T5, and BART illustrated how transformer models could support different text-to-text problems. A common architecture and large-scale pretraining offered a foundation for several applications rather than a separate bespoke model for every task.
Published papers are not always reproducible systems
Only 15% of papers published code in the report’s cited analysis. The report connected this shortfall with accountability and reproducibility, while noting that some industry code depended on proprietary infrastructure. A public paper did not necessarily give another researcher the means to reproduce it.
A paper could describe an important advance without providing the code needed to recreate it. The report’s code-availability analysis made that gap visible. Readers should distinguish learning that a result exists from having the implementation, data, and resources needed to test or extend it.
Evidence you can use
AI progress in the 2020 report
Historical snapshot: October 2020. Dates and populations are specified per row.
The code-availability figure describes the source’s paper population, not every AI project. Parameter count is a model specification, not a score for reasoning or reliability.
GPT-3, T5, and BART illustrated a new generation of transformer models. The report connected them to tasks including translation, summarization, text generation, and code generation.
No. The report highlighted multiple text-to-text uses, including translation and summarization, as well as code generation. Conversational interfaces were not the sole application of the underlying models.
The report described GPT-3 as a 175-billion-parameter model. Parameter count described model scale, not a standalone measure of accuracy or usefulness on every task.
It measured the proportion of papers publishing code in the cited analysis. It did not mean that the remaining papers were necessarily wrong or that every released implementation was reproducible.
The report described code intertwined with proprietary infrastructure and scaling know-how. Publishing a paper did not necessarily expose the full engineering system needed to recreate its results.
They limited the ability of other groups to inspect, reproduce, and build on results. The report connected that concern to a wider concentration of AI talent and compute.
Historical snapshot published October 1, 2020. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
2020 report, PDF page 11Original 2020 report. This edition has no printed slide numbers. References use one-based PDF pages.
2020 report, PDF page 17Original 2020 report. This edition has no printed slide numbers. References use one-based PDF pages.
2020 report, PDF page 24Original 2020 report. This edition has no printed slide numbers. References use one-based PDF pages.
Benaich, Nathan, and Ian Hogarth. “Language models grow more capable and less open.” State of AI Report 2020. Historical report snapshot; web edition prepared 2026-10-11.