All topicsSTATE OF AI REPORT.

Robotics’ GPT-2 moment and the rise of world models

Broader pre-training is helping robots generalize to unfamiliar tasks. This is why I think robotics is approaching its GPT-2 moment, as new models, hardware, and data collection methods come together. And I see this through my investments in companies like Sereact, Wayve, Black Forest Labs and Odyssey.

Questions in this section

Can robots learn an unfamiliar task from a video?Does the 66% result mean robots completed 66% of tasks?How can world models help train robots?Can a robot recover from an error it was not shown in training?Why does a robot need memory?Can simulation predict real robot performance?How widely are robotaxis being used?

Skild's S1 performed unfamiliar tasks from video demonstrations without updating its weights. With 100,000 hours of pre-training, it scored 66% on the company's cumulative per-step success measure, against 9% for a language-prompted baseline trained on the same data and compute.

Robot performance on unseen tasks improves as pre-training data scales.
Robot performance on unseen tasks improves as pre-training data scales.

Robots are learning from fewer demonstrations

Learning from demonstrations could reduce the task-specific training each deployment needs. World models offer another route. Wayve's GAIA-4 turns recorded driving scenes into simulations in which an AI driver's steering and braking change what it sees next. Other road users follow their recorded paths, but the driver can explore the consequences of its own actions.

Odyssey-3 draws on visual pre-training to learn controls for robots, cars, and games. In the company's experiments, robot arms learned from tens of hours of demonstrations and recovered from missed grasps without being shown those recoveries. These early results suggest that physical AI may need fewer demonstrations of every situation it could encounter.

The data must connect observation to action

A video shows what happened, but a robot also needs to learn how to move. Teleoperation records a person controlling the robot directly. Handheld grippers capture demonstrations without requiring the robot at collection time. Wearable cameras can collect human activity at greater scale, but the resulting hand and wrist movements must be translated into the robot’s actions.

Physical Intelligence’s π0.7 addresses a related problem: a large, diverse dataset contains different strategies, mistakes, speeds, and levels of quality. Averaging them together can make behavior worse. The system adds context describing the subtask and how the action was performed. In its laundry experiments, adding lower-quality data improved throughput when it carried this metadata and reduced throughput without it. The quality of the learning signal depends on how the experience is described.

Further reading: slide 56, slide 57.

Context labels make imperfect robot demonstrations useful for training.
Context labels make imperfect robot demonstrations useful for training. Report slide 57.

Reusable skills reduce the demonstrations needed

A kitchen task often combines familiar actions in a new sequence. Penn’s SymSkill learns reusable motions and the conditions under which they can be used, allowing a planner to reorder them. In simulation it reached 85% success across 12 single-step RoboCasa tasks. Separately, a real Franka robot learned from five minutes of play and performed sequences of up to 12 steps.

Those are different experiments, so the 85% figure should not be read as the real robot’s success rate on long tasks. The useful idea is that a robot can recombine practiced movements and recover after disturbances, reducing the need to demonstrate every possible sequence from beginning to end.

Further reading: slide 58.

Memory helps a robot keep track of unfinished work

Long assemblies can look similar at different stages. A camera frame may show the same parts while concealing which operations are already complete. RoboTTT adds adaptive memory to GR00T N1.7 so that earlier observations can inform the next movement.

Across three YAM assembly tasks, it achieved 79% task progress, compared with 42% without memory and 56% with an alternative recurrent memory. Yet on the five-minute Gear Bot assembly it finished only two of ten trials. The baselines finished none. Better progress tracking is valuable, but these results also show how much separates improved component behavior from reliable completion.

Further reading: slide 60.

Robot memory improves task progress, while complete assemblies remain difficult.
Robot memory improves task progress, while complete assemblies remain difficult. Report slide 60.

Simulation is becoming a more useful test environment

Testing every policy on physical hardware is slow and costly. SimFoundry reconstructs interactive scenes from video, then varies objects, layouts, and tasks. Across seven selected tabletop tasks, simulated and real policy scores had a mean correlation of 0.911. That suggests a way to screen policies before using scarce robot time, though the evidence includes optional manual refinement and does not cover arbitrary environments.

Simulation can also accelerate training. FlashSAC trained 4,096 simulated Unitree G1s to climb stairs in four hours on one A100, compared with nearly twenty hours using PPO. The resulting controller transferred to real stairs without further fine-tuning, using randomized physics and an established transfer setup. Faster learning, reusable skills, and memory are complementary ways to make physical deployment less dependent on collecting another large set of task-specific demonstrations.

Further reading: slide 61, slide 62.

Robotaxis are accumulating substantial commercial mileage

Autonomous driving provides evidence of robots operating repeatedly outside a laboratory. Waymo reported 220 million rider-only miles through March 2026, up from 71 million a year earlier. The report records more than 500,000 paid rides a week across 14 US cities. Waymo’s safety hub reported 94% fewer serious-injury crashes than human drivers on the same roads. That comparison comes from the company and applies to its operating locations and measurement method.

Competitors are at different stages and report different measures. Tesla recorded 2.4 million paid robotaxi miles by June 2026. Wayve and Uber offered paid autonomous rides in London with a licensed driver onboard, while Waymo was testing ahead of a planned driverless launch. Baidu’s Apollo Go reported around one million driverless rides in the second quarter, down from 3.2 million in the first, citing operational adjustments on regulatory grounds.

Paid rides, distance traveled, safety outcomes, and whether a driver is onboard each describe a different part of deployment. These measures provide firmer evidence of adoption than a demonstration alone, but they do not establish that the same system can drive everywhere or handle every condition.

Further reading: slide 145.

Waymo reports growing commercial service and rider-only mileage.
Waymo reports growing commercial service and rider-only mileage. Report slide 145.

Evidence you can use

Unseen tasks in Skild’s internal evaluation

S1 release, 2026

Unseen tasks in Skild’s internal evaluation
How the task is specifiedCumulative per-step success
Language-prompted baselineSource9%
Video demonstration in contextSource66%

Company-reported results at 100,000 hours of pre-training, using the same data and compute. The metric is cumulative per-step success across tasks. Human intervention was used to recover from failures during rollouts. These percentages are not end-to-end autonomous task completion rates.

Sources: Skild AI: Introducing S1.

Frequently asked questions

Answers drawn from the report and the sources below.

How can world models help train robots?

Wayve’s GAIA-4 turns recorded driving scenes into simulations in which an AI driver’s steering and braking change what it sees next. Other road users follow their recorded paths, while the driver can explore the consequences of its own actions.

Source: Wayve: GAIA-4.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. Skild AI: Introducing S12026 release. Primary source checked 2026-10-07.
  2. Wayve: GAIA-42026. Retained from the launch essay.
  3. Odyssey: Introducing Odyssey-32026. Retained from the launch essay.
  4. State of AI Report 2026, slide 56: Teaching robots requires data about how to act2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  5. State of AI Report 2026, slide 57: For π0.7, context makes imperfect robot data useful2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 58: A robot turns five minutes of play into reusable skills2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 60: With a longer memory, a robot can improve long-horizon task completion2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 61: Simulation is a bedrock of robotic reality2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 2026, slide 62: A humanoid learns stair climbing in four hours of simulation2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  10. State of AI Report 2026, slide 145: One year on: Waymo tripled to 220M rider-only miles and serves 500k rides a week2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.

Cite this page

Benaich, Nathan. “Robotics’ GPT-2 moment and the rise of world models.” State of AI Report 2026. Published 2026-10-08.