All topicsSTATE OF AI REPORT.

Web knowledge begins to inform physical actions

The 2023 robotics story connected two kinds of progress. Models trained on internet-scale data began to help robots interpret objects and instructions, while specialized learning systems achieved exceptional performance in tightly defined physical tasks.

Questions in this section

What was RT-2?Could RT-2 run quickly enough to control a robot?What did Swift achieve?

RT-2 turns actions into model outputs

RT-2 represented robot actions as tokens and co-fine-tuned vision-language models on robot actions. The aim was to transfer useful knowledge from web data into manipulation: recognizing unfamiliar objects, interpreting new instructions, and choosing an object for its purpose. The robot still depended on a trained control setup and defined action representation.

RT-2 turns actions into model outputs - 2023 report, slide 45
RT-2 turns actions into model outputs. 2023 report, slide 45 (PDF page 45)

Inference speed matters in the physical world

The largest RT-2 model had 55 billion parameters and ran at 1-3 Hz through a multi-TPU cloud service. This detail showed the gap between producing a useful decision and doing so quickly enough for a robot. Capability and deployment cost had to be assessed alongside control frequency.

Swift races against human champions

Swift combined perception, state estimation, and a reinforcement-learning policy trained in simulation. Using onboard sensors and computation, it won several races against three human champions and recorded the fastest time in the reported competition. The achievement was specific to drone racing, rather than evidence of a general-purpose robot.

Swift races against human champions - 2023 report, slide 47
Swift races against human champions. 2023 report, slide 47 (PDF page 47)

Evidence you can use

Robotics in the 2023 report

Historical snapshot: October 2023. Dates and populations are specified per row.

Robotics in the 2023 report
MeasureReported valueDefinition and source
Largest RT-2 model55B parametersLargest vision-language-action model described in the report.2023 report, slide 45 (PDF page 45)
RT-2 control frequency1-3 HzReported inference frequency for the largest model served through a multi-TPU cloud system.2023 report, slide 45 (PDF page 45)
Swift human opponents3 championsHuman champions in the reported drone-racing comparison; Swift won several races.2023 report, slide 47 (PDF page 47)

RT-2 specifications and Swift racing outcomes describe different systems. Racing performance does not establish general robotics capability, and inference frequency is not a success rate for manipulation tasks.

Frequently asked questions

What was RT-2?

RT-2 was a vision-language-action approach that represented robot actions as tokens and trained jointly on web-derived knowledge and robot data, allowing semantic knowledge to inform manipulation.

Source: 2023 report, slide 45 (PDF page 45).

Sources and dates

Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2023 report, slide 45 (PDF page 45)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  2. 2023 report, slide 47 (PDF page 47)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  3. State of AI Report 2023: online slides
  4. Nathan Benaich: The State of AI Report 2023Air Street Press, October 12, 2023.
  5. Welcome to State of AI Report 2023Original website launch post, October 12, 2023.

Cite this page

Benaich, Nathan. “Web knowledge begins to inform physical actions.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.