As models and the scaffolding around them improve together (and build themselves), I expect many of today's weaknesses to be learned away. We can also get more from capabilities that are already available.
Different tasks benefit from different agent setups
An agent harness is the software that gives a model tools, supplies context, stores state, and handles feedback. The harness can determine whether the model gets the right evidence and whether a failed attempt leads to a useful retry. This is why comparing models inside different agent products can confound model capability with the environment around it.
Researchers at Meta, Duke, and UC Davis tested a further step: evolve two specialized harnesses and choose between them for each new problem. With Gemini 3 Flash held fixed, the router reached 62% on the math evaluation versus 46% for Meta-Harness. The two branches together could potentially solve 67%, leaving room to improve the routing decision. Coding results also improved with a fixed Sonnet 4.5 model. A setup that wins on average can still be the wrong one for a particular task.
Further reading: slide 7, slide 8.

Long context needs a way to find the right information
Putting more text in a prompt does not ensure that a model will use it well. MIT’s Recursive Language Models keep the input in a code workspace. The model can inspect it, select relevant pieces, delegate smaller questions to further model calls, and assemble the results. Documents become something the agent can work on, rather than something it must absorb in one pass.
The approach lets a fixed model tackle inputs that are too large for a single call. Training on short tasks also transferred to longer inputs and new domains when they shared a useful way of decomposing the problem. The benefit depends on that decomposition, and additional calls can increase cost. Choosing what to read is part of the agent’s work.
Further reading: slide 9.
Skills and memory preserve useful experience
Skills package reusable instructions and code. Memory preserves information for later tasks. Both allow an agent to benefit from prior work without changing the underlying model’s weights. A tested procedure for checking a spreadsheet or a record of a project decision can save the next run from rediscovering the same information.
The production loop in the report explains how to improve these systems deliberately. Record the task, tools, outcome, and corrections. Turn recurring failures into repeatable tests, then decide whether to change context, memory, routing, a tool, or the model. Recording whether each action succeeded gives the team evidence for the next change. Accumulating transcripts without evaluating what worked does not establish learning.
Further reading: slide 10, slide 114.
Better models can need fewer behavioral instructions
Some harness features compensate for weaknesses that a later model no longer has. The report describes Claude Code removing 80% of its system prompt for advanced models without measurable loss on coding evaluations. Examples intended to teach tool use can also constrain exploration when the model already understands the tools.
Harnesses need to be retested as models improve. Their design still needs clear tool interfaces, reliable state, observable actions, and controlled execution environments. After a model upgrade, retest the old planning rules and reminders as well as the model itself. The software surrounding an agent should make work easier to perform and verify.
Further reading: slide 14.
A capable agent can still make poor decisions over months
Long-running work requires an agent to manage resources, judge risk, and learn from previous decisions. KellyBench tests this in a simulated English Premier League betting market. Agents receive a season’s information sequentially and start with £100,000 to build models and size bets. All 12 tested models lost money on average over five runs, and six went bankrupt at least once.
E-CommerceBench gives 18 models CNY 100,000 to operate up to four stores for 365 simulated days. GPT-5.6 Sol produced the highest reported average year-end assets, CNY 1.43 million, while directing 18.48% of order spending to fraudulent suppliers. Opus 4.7 directed 0.12% to fraudulent suppliers. Sixteen of the 18 models showed no clear improvement in bargaining on repeat purchases.
These are simulated markets with particular rules and incentives. They do not measure actual business profitability. They do show that high returns on one measure can coexist with weak supplier screening, and that accumulating experience does not guarantee better decisions. Assessing an autonomous business agent requires examining losses, fraud exposure, and learning as well as its final balance.
Further reading: slide 44, slide 45.
Agent adoption extends beyond developers
What’s exciting is how far this extends beyond developers. OpenAI's study of agent adoption finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.

Models are helping build their successors
AI is already helping build better AI. According to Anthropic's internal index, Claude led 26% of measured model R&D work in August, up from under 1% in February. Researchers set the tasks and supervise execution, and Anthropic reports that this is speeding up development. This is a branch of AI research that I am most excited about.

For many, Karpathy's autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.
What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.