All topicsSTATE OF AI REPORT. 2021

Transformers become a general-purpose architecture

The 2021 report documented a broadening of transformers beyond language. Images could be treated as sequences of patches, while paired text and images supported reusable representations. Strong benchmark performance still left important weaknesses, including a tendency to reproduce falsehoods.

Questions in this section

Why were transformers important beyond language in 2021?Did larger language models always give truer answers?What was the central idea of a Vision Transformer?Did transformer success in vision depend only on the architecture?How did CLIP connect text and images?Why did CLIP’s zero-shot results matter?How did the strongest reported model compare with humans on TruthfulQA?What was the role of the world model in DreamerV2?

Vision transformers change the input, not the core idea

A Vision Transformer splits an image into patches and processes them as a sequence. The report described strong ImageNet results as both model size and training data increased. Hybrid architectures combining attention and convolutions remained competitive, so this was not the end of convolutional networks.

Vision transformers change the input, not the core idea - 2021 report, slide 11
Vision transformers change the input, not the core idea. 2021 report, slide 11 (PDF page 11)

Text and images become a shared training signal

CLIP learned from 400 million text-image pairs and could classify across several datasets without task-specific fine-tuning. The result showed the value of connecting language with visual representations, while performance still depended on the evaluation and prompts.

Text and images become a shared training signal - 2021 report, slide 38
Text and images become a shared training signal. 2021 report, slide 38 (PDF page 38)

Larger models are not automatically more truthful

On TruthfulQA, the best model in the reported comparison was truthful on 58% of questions versus 94% for humans. The questions were designed to expose common falsehoods, so the result concerned a difficult targeted benchmark rather than all everyday answers.

Larger models are not automatically more truthful - 2021 report, slide 44
Larger models are not automatically more truthful. 2021 report, slide 44 (PDF page 44)

World models let agents rehearse before acting

DreamerV2 learned a compact model of an Atari environment from pixels, then used that model to learn behavior. The report highlighted strong performance on a 55-task Atari benchmark using a single GPU. The important shift was where learning occurred: the agent could explore possible behavior inside a learned representation of the environment, reducing its dependence on expensive interactions with the game.

World models let agents rehearse before acting - 2021 report, slide 29
World models let agents rehearse before acting. 2021 report, slide 29 (PDF page 29)

Evidence you can use

AI progress in the 2021 report

Historical snapshot: October 2021. Dates and populations are specified per row.

AI progress in the 2021 report
MeasureReported valueDefinition and source
CLIP training pairs400MText-image pairs used for pretraining.2021 report, slide 38 (PDF page 38)
Best model truthfulness58%TruthfulQA comparison reported in the 2021 deck.2021 report, slide 44 (PDF page 44)
Human baseline truthfulness94%Human comparison on the same benchmark.2021 report, slide 44 (PDF page 44)

Training size is not a quality score. TruthfulQA was designed to expose particular failure modes; its rates should not be treated as universal accuracy measures.

Frequently asked questions

Sources and dates

Historical snapshot published October 12, 2021. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2021 report, slide 11 (PDF page 11)Original 2021 report. Printed slide numbers match PDF pages.
  2. 2021 report, slide 29 (PDF page 29)Original 2021 report. Printed slide numbers match PDF pages.
  3. 2021 report, slide 38 (PDF page 38)Original 2021 report. Printed slide numbers match PDF pages.
  4. 2021 report, slide 44 (PDF page 44)Original 2021 report. Printed slide numbers match PDF pages.
  5. State of AI Report 2021: online slides
  6. Air Street Press launch essayOctober 12, 2021.
  7. Welcome to State of AI Report 2021Original website launch post, October 12, 2021.

Cite this page

Benaich, Nathan, and Ian Hogarth. “Transformers become a general-purpose architecture.” State of AI Report 2021. Historical report snapshot; web edition prepared 2026-10-11.