All topicsSTATE OF AI REPORT.

The frontier is now a three-lab race

The frontier is now a three-lab race between Anthropic, OpenAI, and Google. Anthropic leads on Artificial Analysis's Intelligence Index, while Google leads on Arena's ranking of the answers people prefer, for now.

Questions in this section

Is AI already helping researchers build better models?Who leads the AI frontier in 2026?Has autonomous recursive self-improvement been demonstrated?What does scientific taste mean for AI?Does a high benchmark score mean the whole task is complete?What does METR’s time horizon measure?Can AI generate interactive video in real time?Can AI tutoring improve student learning?
The frontier is now a three-lab race, with different leaders on the Intelligence Index and Arena.
The frontier is now a three-lab race, with different leaders on the Intelligence Index and Arena.

Last year, we described reasoning becoming useful at scale with OpenAI’s o-series of models. I have lived through the improvements since then, as the amount of useful work I can iterate on and increasingly delegate to agents keeps growing.

My view is that much of the gap between people getting substantial value from AI and those getting little comes down to knowing how to use it and how to set it up. Put another way, it’s a “human skill issue,” not an “AI technology issue.” I really believe that we need a Genius Bar for AI, where people can get hands-on help choosing tools and setting them up to help them in their everyday tasks.

Benchmarks are approaching their ceilings

Across mathematics, scientific reasoning, and coding, benchmarks intended to challenge models for years are approaching their ceilings within months. We need new evaluations that reveal where models still separate and what they can reliably do.

Benchmarks featured in last year's report have rapidly approached their ceilings.
Benchmarks featured in last year's report have rapidly approached their ceilings.

AI is accelerating the work of building AI

AI is already helping build better AI. According to Anthropic's internal index, Claude led 26% of measured model R&D work in August, up from under 1% in February. Researchers set the tasks and supervise execution, and Anthropic reports that this is speeding up development. This is a branch of AI research that I am most excited about.

Claude leads a growing share of Anthropic's measured model R&D, under human supervision.
Claude leads a growing share of Anthropic's measured model R&D, under human supervision.

For many, Karpathy's autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.

What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.

Open models are changing who builds on AI

Our analysis with Zeta Alpha tracks AI papers on arXiv that mention 21 model families. Chinese open-weight families gained ground sharply between 2024 and 2026, and Qwen overtook Llama. Closed US models still accounted for 42% of mentions in the 2026 snapshot.

Paper mentions reveal which models researchers choose to study and build on. They do not measure commercial revenue or establish a universal capability ranking. A model used widely in research can become the starting point for new methods, evaluations, and specialist applications. The research ecosystem can shift even while American labs retain leading closed models.

Further reading: slide 6.

Epoch AI found flaws in nine of 15 benchmarks

A high score is only useful if the test measures what it claims to measure. Epoch AI found substantive flaws in nine of its first 15 benchmark reviews. It examined tasks, grading, prompts, tools, and resource limits, finding problems such as broken scoring and exploitable environments. Four benchmarks met its minimum standards with caveats. Two could not be judged with the available information.

Even a sound evaluation answers a particular question. FrontierSWE gives agents 20 hours to build software in a common harness. MirrorCode asks them to reconstruct programs from documentation and a runnable reference, with up to seven days and ten billion tokens. They produce different leaders. The task, completion criteria, and spending allowance belong beside the model’s name whenever we compare results.

Further reading: slide 36, slide 41.

Progress on a task is different from finishing it

On FrontierChallenge’s 97 scientific workflows, GPT-5.6 Sol earned 87.9 out of 100 for satisfying individual requirements but fully completed only 20.6% of tasks. On OSWorld 2.0’s computer workflows, Opus 5 earned 77.7% partial credit and completed 44.3%. An otherwise correct expense claim can remain unfinished because the agent misses an approval or never submits it.

Task duration needs similar care. METR’s 50% time horizon estimates the human-expert time associated with tasks an agent can complete half the time. It is not a measurement of how long an agent can safely operate unattended. With only five tasks taking humans more than 16 hours, METR warns that estimates at the upper end are unreliable. Longer and better-scored evaluations are necessary to tell whether capability is becoming dependable.

Further reading: slide 42, slide 43.

Partial benchmark scores can substantially exceed complete-task success.
Partial benchmark scores can substantially exceed complete-task success. Report slide 42.

Research engineering is advancing faster than research judgment

In a shadow evaluation, agents received the central questions from unpublished NeurIPS submissions and the original authors reviewed their work. Each run had six days, GPUs, and $3,000 of API credit. Opus 4.8 completed the engineering without human help, but its two papers received rejection scores of 2/6 and 1/6. The recurring weaknesses included uncreative fixes, ineffective backtracking, and a poor sense of what would merit publication.

Agents can make experiments easier to implement while people still choose the research questions and judge the results. Demonstrating sustained autonomous research requires evidence that the system can make those choices too.

Further reading: slide 17, slide 23.

Video generation is becoming fast enough for interactive streams

Video generation usually returns a completed clip after a prompt. fal adapted MiniMax’s open-weight H3 video and audio model into H3 Max, combining additional training data with a serving engine and optimized GPU kernels. The company reports producing a five-second clip in under three seconds, with about 35 times the throughput of the official H3 endpoint.

Its Director mode carries 39 frames into the next segment and remembers previous prompts. Removing the overlapping frames lets the segments play continuously. On fal.live, viewers vote on what happens next in a 24-frame-per-second stream. Their prompts affect the following segment, so the interaction still has a delay. Faster generation creates room for audience-directed video and other experiences in which the scene changes while someone watches.

Further reading: slide 48.

Continuous video generation lets viewers steer the next segment.
Continuous video generation lets viewers steer the next segment. Report slide 48.

AI tutoring needs to help students solve problems themselves

A tutor can give the right answer while leaving the student unable to solve the next problem. A trial in Sierra Leone tested a more structured use of AI. Across 48 classrooms in 12 schools, 1,763 students were randomized to teacher-led Gemini activities or standard instruction. After eight weeks, the AI program improved math scores by 0.258 standard deviations, with a 95% confidence interval of 0.027 to 0.488.

The result applies to that classroom program, including its teachers and activities. It does not establish the effect of unrestricted chatbot use or long-term learning. TutorMoments and MathTutorBench also distinguish answering correctly from teaching effectively. Useful tutoring needs to decide when to give a hint, when to ask the student to explain, and when to let them work through a mistake.

Further reading: slide 105.

A classroom trial measured learning from teacher-led AI activities.
A classroom trial measured learning from teacher-led AI activities. Report slide 105.

Evidence you can use

Claude’s share of measured model R&D

February-August 2026

Claude’s share of measured model R&D
PeriodShare led by Claude
February 2026SourceUnder 1%
August 2026Source26%

Anthropic’s internal index measures work inside one lab, under human supervision. It does not measure the share of all AI research automated or establish that AI can run a research agenda on its own.

Sources: Anthropic: measuring the pace of AI development.

Frequently asked questions

Answers drawn from the report and the sources below.

Who leads the AI frontier in 2026?

The report describes a three-lab race between Anthropic, OpenAI, and Google. Its snapshot places Anthropic first on Artificial Analysis’s Intelligence Index and Google first on Arena’s ranking of answers people prefer. The leader depends on the measure, and these rankings change.

Source: State of AI Report 2026.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. Anthropic: measuring the pace of AI developmentFebruary-August 2026. Retained from the launch essay.
  2. Anthropic: recursive self-improvement2026. Retained from the launch essay.
  3. Karpathy: autoresearchRepository, changing over time. Retained from the launch essay.
  4. Nathan Benaich: Can AI learn scientific taste?2026. Author analysis.
  5. State of AI Report 2026, slide 6: Chinese open-weight models overtook American ones in AI research papers in 20262026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 17: Can agents produce a top-tier research paper? No, but they can do its engineering.2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 23: As task benchmarks saturate, RSI evidence is moving inside the labs2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 36: But who benchmarks the benchmarks?2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 2026, slide 41: Long-horizon coding rankings change with the task and the evaluation budget2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  10. State of AI Report 2026, slide 42: High scores can hide unfinished scientific analyses and desk work2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  11. State of AI Report 2026, slide 43: METR needs harder tasks to reliably measure the strongest models2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  12. State of AI Report 2026, slide 48: Generative video goes real time and lets a streamer steer it!2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  13. State of AI Report 2026, slide 105: AI in education: the best tutor is not a helpful assistant2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  14. State of AI Report 20262026 report snapshot. Report PDF. See individual slide references for the expanded analysis.

Cite this page

Benaich, Nathan. “The frontier is now a three-lab race.” State of AI Report 2026. Published 2026-10-08.