
How to read this finding
Video hours are training inputs, not successful task hours. SayCan’s rates concern its tested instructions, skills, and environment; they are not general-purpose robotics reliability.
It first learned to infer actions from labeled videos, then used that model to label a much larger video collection for training. Further imitation and reinforcement learning developed more complex behavior.

Video hours are training inputs, not successful task hours. SayCan’s rates concern its tested instructions, skills, and environment; they are not general-purpose robotics reliability.
Evidence you can use
Historical snapshot: October 2022. Dates and populations are specified per row.
| Measure | Reported value | Definition and source |
|---|---|---|
| VPT action-labeled video | 2,000 hours | Initial video with mouse and keyboard labels.2022 report, slide 21 (PDF page 21) |
| VPT additional video | 70,000 hours | Video labeled using the learned inverse-dynamics model.2022 report, slide 21 (PDF page 21) |
| SayCan execution success | 74% | Reported evaluation across 101 instructions, with 84% planning success.2022 report, slide 37 (PDF page 37) |
Video hours are training inputs, not successful task hours. SayCan’s rates concern its tested instructions, skills, and environment; they are not general-purpose robotics reliability.
Historical snapshot published October 11, 2022. This web edition was prepared on 2026-10-11 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan, and Ian Hogarth. “How did VPT learn to play Minecraft?” State of AI Report 2022. Historical report snapshot; web edition prepared 2026-10-11.