RT-2 turns actions into model outputs
RT-2 represented robot actions as tokens and co-fine-tuned vision-language models on robot actions. The aim was to transfer useful knowledge from web data into manipulation: recognizing unfamiliar objects, interpreting new instructions, and choosing an object for its purpose. The robot still depended on a trained control setup and defined action representation.

Inference speed matters in the physical world
The largest RT-2 model had 55 billion parameters and ran at 1-3 Hz through a multi-TPU cloud service. This detail showed the gap between producing a useful decision and doing so quickly enough for a robot. Capability and deployment cost had to be assessed alongside control frequency.
Swift races against human champions
Swift combined perception, state estimation, and a reinforcement-learning policy trained in simulation. Using onboard sensors and computation, it won several races against three human champions and recorded the fastest time in the reported competition. The achievement was specific to drone racing, rather than evidence of a general-purpose robot.
