Models spend more compute on an answer
OpenAI o1 used reinforcement learning to improve step-by-step reasoning. The report highlighted large gains on competition mathematics, alongside a practical tradeoff: o1-preview was slower and more expensive than GPT-4o and lacked some of its features. More time spent reasoning created another way to improve performance, but did not make it the best model for every task.

Open models reach the frontier
Llama 3.1 405B competed with GPT-4o and Claude 3.5 Sonnet across several reasoning, mathematics, multilingual, and long-context benchmarks. Meta trained the family on 15 trillion tokens and followed it with multimodal and on-device models. Availability of weights expanded what others could build and inspect, while the model’s license still mattered.

Strong reasoning remains uneven
The first tests of o1 showed striking results on some complex mathematics and science problems, alongside weaknesses in spatial reasoning and games such as chess. The report treated these as evidence of uneven capabilities. A high benchmark score did not establish reliable reasoning across unfamiliar situations.
