Autonomous research and scientific judgment
For many, Karpathy's autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.
What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.
Frequently asked questions
Answers drawn from the report and the sources below.
The report describes a three-lab race between Anthropic, OpenAI, and Google. Its snapshot places Anthropic first on Artificial Analysis’s Intelligence Index and Google first on Arena’s ranking of answers people prefer. The leader depends on the measure, and these rankings change.
Source: State of AI Report 2026.
The work described in the report automates parts of the research process. Karpathy’s autoresearch, for example, lets an agent edit training code and run short experiments. That falls short of demonstrating sustained, fully autonomous recursive improvement.
Source: Karpathy: autoresearch · Anthropic: recursive self-improvement.
It means choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. My argument is that learning this judgment may require the alternatives, failures, and decisions that published papers leave out.
Source: Nathan Benaich: Can AI learn scientific taste?.