The 2022 report described early routes from generating text to performing actions. Models learned from videos, used search interfaces, and proposed robot actions. Each route needed a connection between the model’s outputs and what an environment or tool could actually do.
Video pretraining supplies experience for an agent
OpenAI’s VPT used 2,000 hours of action-labeled video to learn how to infer actions, then labeled 70,000 hours of Minecraft video. Further training produced behaviors that were difficult to learn through reinforcement learning alone. Minecraft provided a defined test environment, not evidence of unrestricted computer autonomy.
Tool use connects language models to external information
WebGPT learned to use a search interface from human demonstrations, allowing it to produce answers grounded in retrieved references. The report also described early commercial work on interacting with websites and software. A tool connection expanded what a model could access but did not guarantee that every answer or action was correct.
PaLM-SayCan combined language-model suggestions with estimates of which robot skills could succeed in the current environment. In the reported evaluation, planning succeeded on 84% of instructions and execution on 74%. The distinction showed why a plausible plan was not the same as completed physical work.
Choosing a useful action was not the same as executing it
SayCan separated a language model’s suggestion from an estimate of which robot skills were feasible. Its different planning and execution success rates made that distinction measurable. An agent needed both a sensible sequence of actions and the practical ability to carry them out in its environment.
Evidence you can use
AI agents and robotics in the 2022 report
Historical snapshot: October 2022. Dates and populations are specified per row.
Video hours are training inputs, not successful task hours. SayCan’s rates concern its tested instructions, skills, and environment; they are not general-purpose robotics reliability.
It first learned to infer actions from labeled videos, then used that model to label a much larger video collection for training. Further imitation and reinforcement learning developed more complex behavior.
A language model could suggest a sensible plan without knowing what the robot could execute. Skill-success estimates helped select actions feasible in the actual environment.
Those videos paired what appeared on screen with mouse and keyboard actions. The report described using them to train an inverse-dynamics model that could infer actions in additional video.
The report cited 70,000 hours labeled using the learned inverse-dynamics model, after an initial 2,000 hours of action-labeled video. The two datasets served different roles.
WebGPT used web search and references when producing answers. The report presented this as a way for language models to interact with external information rather than rely only on stored training knowledge.
No. The report described promising results from tool use, not a universal guarantee. The value depended on how the model chose tools, interpreted outputs, and incorporated evidence into its answer.
Across 101 instructions, the report cited 84% planning success and 74% execution success. A correct plan could still fail when translated into physical actions.
They described the report’s evaluation across 101 instructions from seven language-instruction types. They were not general success rates for every robot, environment, or user request.
Historical snapshot published October 11, 2022. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “Models learn to use tools and act in environments.” State of AI Report 2022. Historical report snapshot; web edition prepared 2026-10-11.