By Nathan Benaich (Air Street Capital) · 2025 report
In 2025, progress increasingly came from what a model did before giving its answer: taking more steps, checking its work, and trying alternatives. OpenAI retained a narrow frontier lead, but DeepSeek, Qwen, and Kimi made open-weight models credible alternatives on reasoning and coding.
OpenAI’s o1 helped establish inference-time computation as a way to improve difficult answers. Instead of relying only on a larger training run, a system could spend more computation on a particular problem. Reinforcement learning with verifiable rewards then gave models feedback from answers that could be checked automatically, such as a mathematical result or a passing program test.
The report found Qwen accounting for more than 40% of new monthly model derivatives on Hugging Face, while Llama’s share had fallen to about 15%. That measures what developers were building on, rather than downloads, revenue, or all AI usage. Together with stronger DeepSeek and Kimi models, it showed China becoming a major source of capable, adaptable models.
Olympiad-level mathematics and advances in formal theorem proving showed substantial capability gains. But the report also documented failures from distracting facts and small changes to problem wording. Studies disagreed on whether reinforcement learning created new reasoning abilities or made existing successful paths easier to sample. A strong score on familiar tasks did not settle that question.
These shares refer to new monthly model derivatives on Hugging Face. They do not measure all model deployments, revenue, downloads, or the share of AI research. Approximate values are preserved as reported.
OpenAI’s GPT-5 variants still led across the independent leaderboards discussed in the report. The gap had narrowed, with Chinese open-weight models and other US labs close on reasoning and coding. Leadership depended on the benchmark.
Models could improve difficult answers by spending more computation at inference time. Training with automatically verifiable rewards, especially in mathematics and programming, helped make those reasoning paths more useful.
No. The reported shift concerns new monthly model derivatives on Hugging Face: more than 40% for Qwen and about 15% for Llama at the snapshot. It is evidence of developer adoption in that ecosystem.
Not consistently. The report describes sensitivity to irrelevant facts, changed numbers, and different reasoning formats. Mathematical benchmark progress did not eliminate those failure modes.
Historical snapshot published October 9, 2025. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.
2025 report, slide 12 (PDF page 13)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 18 (PDF page 19)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 21 (PDF page 22)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 22 (PDF page 23)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 30 (PDF page 31)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 32 (PDF page 33)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 33 (PDF page 34)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.
2025 report, slide 45 (PDF page 46)Original 2025 report. Slide numbers printed in the deck are one lower than PDF page numbers because the cover is unnumbered.