
How to read this finding
RT-2 specifications and Swift racing outcomes describe different systems. Racing performance does not establish general robotics capability, and inference frequency is not a success rate for manipulation tasks.
RT-2 was a vision-language-action approach that represented robot actions as tokens and trained jointly on web-derived knowledge and robot data, allowing semantic knowledge to inform manipulation.

RT-2 specifications and Swift racing outcomes describe different systems. Racing performance does not establish general robotics capability, and inference frequency is not a success rate for manipulation tasks.
Evidence you can use
Historical snapshot: October 2023. Dates and populations are specified per row.
| Measure | Reported value | Definition and source |
|---|---|---|
| Largest RT-2 model | 55B parameters | Largest vision-language-action model described in the report.2023 report, slide 45 (PDF page 45) |
| RT-2 control frequency | 1-3 Hz | Reported inference frequency for the largest model served through a multi-TPU cloud system.2023 report, slide 45 (PDF page 45) |
| Swift human opponents | 3 champions | Human champions in the reported drone-racing comparison; Swift won several races.2023 report, slide 47 (PDF page 47) |
RT-2 specifications and Swift racing outcomes describe different systems. Racing performance does not establish general robotics capability, and inference frequency is not a success rate for manipulation tasks.
Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
Benaich, Nathan. “What was RT-2?.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.