All topicsSTATE OF AI REPORT.

GPT-4 leads as open models gather momentum

GPT-4 was the reference point for language-model capability in the 2023 report. At the same time, Llama 2 made capable models more accessible, and experiments with smaller models showed how much carefully chosen training data could matter. Progress came with growing uncertainty about what closed labs were actually building.

Questions in this section

What did Llama 2 change in 2023?Did GPT-4 stop hallucinating?Could small models rival larger ones?

GPT-4 sets the benchmark

The report described GPT-4 as a substantial improvement over its predecessors, including stronger performance on reasoning and knowledge tasks. Reinforcement learning from human feedback improved some behaviors, but hallucinations remained. Better answers did not remove the need to examine how a model was evaluated or where it still failed.

GPT-4 sets the benchmark - 2023 report, slide 12
GPT-4 sets the benchmark. 2023 report, slide 12 (PDF page 12)

Llama 2 gives builders another route

Meta’s Llama 2 used a two-trillion-token pretraining corpus and additional instruction tuning and human feedback. The 70B model was competitive with ChatGPT on many tasks but lagged in coding, where a specialized Code Llama variant performed better. The release permitted broad commercial use subject to license conditions, rather than unrestricted use by every company.

Llama 2 gives builders another route - 2023 report, slide 19
Llama 2 gives builders another route. 2023 report, slide 19 (PDF page 19)

Small models benefit from carefully selected data

Microsoft’s TinyStories and phi work explored whether narrow, curated datasets could produce useful capabilities in much smaller models. A 28-million-parameter story model compared favorably with a much larger model under a GPT-4-based evaluation. These were task-specific results; they did not establish that a small model could replace a frontier model across all uses.

Small models benefit from carefully selected data - 2023 report, slide 26
Small models benefit from carefully selected data. 2023 report, slide 26 (PDF page 26)

Evidence you can use

AI progress in the 2023 report

Historical snapshot: October 2023. Dates and populations are specified per row.

AI progress in the 2023 report
MeasureReported valueDefinition and source
Llama 2 pretraining corpus2T tokensCorpus size reported for Llama 2, a 40% increase over its predecessor.2023 report, slide 19 (PDF page 19)
Llama 2 largest model discussed70B parametersModel compared with ChatGPT across tasks in the report.2023 report, slide 19 (PDF page 19)
phi-1 model size1.3B parametersSmall code model trained on curated code and synthetic educational material.2023 report, slide 26 (PDF page 26)

Model size and corpus size are specifications, not quality scores. Comparisons in the report used different tasks and evaluators. Llama’s availability was subject to its license; small-model findings were limited to the evaluated settings.

Frequently asked questions

Sources and dates

Historical snapshot published October 12, 2023. This web edition was prepared on October 10, 2026 from the online deck and original launch posts. Findings and forecasts retain their original time frame.

  1. 2023 report, slide 12 (PDF page 12)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  2. 2023 report, slide 19 (PDF page 19)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  3. 2023 report, slide 26 (PDF page 26)Original 2023 report. Printed slide numbers match PDF page numbers in this edition.
  4. State of AI Report 2023: online slides
  5. Nathan Benaich: The State of AI Report 2023Air Street Press, October 12, 2023.
  6. Welcome to State of AI Report 2023Original website launch post, October 12, 2023.

Cite this page

Benaich, Nathan. “GPT-4 leads as open models gather momentum.” State of AI Report 2023. Historical report snapshot; web edition prepared 2026-10-10.