All topicsSTATE OF AI REPORT.

Reasoning opens another route to scale

By October 2024, the leading language models were looking increasingly alike on familiar benchmarks. OpenAI o1 changed the question from how much compute to spend training a model to how much it should spend working through a problem. At the same time, Meta showed that an openly available model could compete near the proprietary frontier.

Questions in this section

What changed with OpenAI o1 in 2024?Did open models catch up in 2024?Did o1 solve reasoning?

Models spend more compute on an answer

OpenAI o1 used reinforcement learning to improve step-by-step reasoning. The report highlighted large gains on competition mathematics, alongside a practical tradeoff: o1-preview was slower and more expensive than GPT-4o and lacked some of its features. More time spent reasoning created another way to improve performance, but did not make it the best model for every task.

Models spend more compute on an answer - 2024 report, slide 13
Models spend more compute on an answer. 2024 report, slide 13 (PDF page 14)

Open models reach the frontier

Llama 3.1 405B competed with GPT-4o and Claude 3.5 Sonnet across several reasoning, mathematics, multilingual, and long-context benchmarks. Meta trained the family on 15 trillion tokens and followed it with multimodal and on-device models. Availability of weights expanded what others could build and inspect, while the model’s license still mattered.

Open models reach the frontier - 2024 report, slide 15
Open models reach the frontier. 2024 report, slide 15 (PDF page 16)

Strong reasoning remains uneven

The first tests of o1 showed striking results on some complex mathematics and science problems, alongside weaknesses in spatial reasoning and games such as chess. The report treated these as evidence of uneven capabilities. A high benchmark score did not establish reliable reasoning across unfamiliar situations.

Strong reasoning remains uneven - 2024 report, slide 14
Strong reasoning remains uneven. 2024 report, slide 14 (PDF page 15)

Evidence you can use

AI progress in the 2024 report

Historical snapshot: October 2024. Dates and populations are specified per row.

AI progress in the 2024 report
MeasureReported valueDefinition and source
Llama training corpus15T tokensTraining corpus size reported for the Llama 3 family.2024 report, slide 15 (PDF page 16)
Llama 3.1 largest model405B parametersThe largest model discussed in the Llama 3.1 release.2024 report, slide 15 (PDF page 16)
o1-preview output price$60 per 1M tokensAPI list price in the October 2024 report; a historical price, not a current quote.2024 report, slide 13 (PDF page 14)

These are reported model specifications and launch-era prices. Benchmark comparisons depend on task, inference budget, and evaluation setup; they are not a universal ranking of intelligence.

Frequently asked questions

Sources and dates

Historical snapshot published October 10, 2024. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2024 report, slide 13 (PDF page 14)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  2. 2024 report, slide 14 (PDF page 15)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  3. 2024 report, slide 15 (PDF page 16)Original 2024 report. Printed slide numbers are one lower than PDF page numbers because the cover is unnumbered.
  4. State of AI Report 2024: online slides
  5. Nathan Benaich: The State of AI Report 2024Air Street Press, October 10, 2024.
  6. Welcome to State of AI Report 2024Original website launch post, October 10, 2024.

Cite this page

Benaich, Nathan. “Reasoning opens another route to scale.” State of AI Report 2024. Historical report snapshot; web edition prepared 2026-10-10.