Transformers become a general-purpose architecture
By Nathan Benaich and Ian Hogarth · 2021 report
The 2021 report documented a broadening of transformers beyond language. Images could be treated as sequences of patches, while paired text and images supported reusable representations. Strong benchmark performance still left important weaknesses, including a tendency to reproduce falsehoods.
Vision transformers change the input, not the core idea
A Vision Transformer splits an image into patches and processes them as a sequence. The report described strong ImageNet results as both model size and training data increased. Hybrid architectures combining attention and convolutions remained competitive, so this was not the end of convolutional networks.
CLIP learned from 400 million text-image pairs and could classify across several datasets without task-specific fine-tuning. The result showed the value of connecting language with visual representations, while performance still depended on the evaluation and prompts.
On TruthfulQA, the best model in the reported comparison was truthful on 58% of questions versus 94% for humans. The questions were designed to expose common falsehoods, so the result concerned a difficult targeted benchmark rather than all everyday answers.
DreamerV2 learned a compact model of an Atari environment from pixels, then used that model to learn behavior. The report highlighted strong performance on a 55-task Atari benchmark using a single GPU. The important shift was where learning occurred: the agent could explore possible behavior inside a learned representation of the environment, reducing its dependence on expensive interactions with the game.
Training size is not a quality score. TruthfulQA was designed to expose particular failure modes; its rates should not be treated as universal accuracy measures.
Vision transformers processed image patches as sequences, while models such as CLIP learned from paired text and images. This made the architecture useful across different types of data.
The system divided an image into patches and treated them as a sequence for a transformer. The report presented this as a way to apply an architecture prominent in language to visual recognition.
No. The report emphasized the benefits of scaling and large pretraining datasets. The results reflected a combination of architecture, data, and training, rather than a change in model design alone.
It learned from paired images and text, giving the system a shared basis for matching visual content with language descriptions. The report cited 400 million training pairs.
They showed that language descriptions could guide classification without a separate supervised training run for every evaluated dataset. The finding concerned the reported tasks, rather than unrestricted visual understanding.
The report cited 58% truthfulness for the best-performing model and 94% for the human baseline. Those figures belonged to a benchmark designed to probe false or misleading answers.
It learned a representation of the environment from pixels so the agent could learn behavior within that representation. The report highlighted its results on a 55-task Atari benchmark using one GPU, illustrating a route to more computationally efficient reinforcement learning.
Historical snapshot published October 12, 2021. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “Transformers become a general-purpose architecture.” State of AI Report 2021. Historical report snapshot; web edition prepared 2026-10-11.