All topicsSTATE OF AI REPORT.

Better tools and context make agents more capable

In a controlled coding study, changing the tools, context, and feedback available to GLM-5.1 lifted success from 52.5% to 65.5% on 100 SWE-bench Verified tasks, without changing its weights. Yes, this is an “older model,” but the principle holds.

Questions in this section

Can an AI agent improve without changing its model?What is an agent harness?Does the study show a 13% improvement?Are AI agents only useful for developers?Can memory improve an agent without retraining it?Should an agent always use the same harness?Can an AI agent reliably run a business for a year?

As models and the scaffolding around them improve together (and build themselves), I expect many of today's weaknesses to be learned away. We can also get more from capabilities that are already available.

Different tasks benefit from different agent setups

An agent harness is the software that gives a model tools, supplies context, stores state, and handles feedback. The harness can determine whether the model gets the right evidence and whether a failed attempt leads to a useful retry. This is why comparing models inside different agent products can confound model capability with the environment around it.

Researchers at Meta, Duke, and UC Davis tested a further step: evolve two specialized harnesses and choose between them for each new problem. With Gemini 3 Flash held fixed, the router reached 62% on the math evaluation versus 46% for Meta-Harness. The two branches together could potentially solve 67%, leaving room to improve the routing decision. Coding results also improved with a fixed Sonnet 4.5 model. A setup that wins on average can still be the wrong one for a particular task.

Further reading: slide 7, slide 8.

Specialized harnesses solve different problems, making routing useful.
Specialized harnesses solve different problems, making routing useful. Report slide 8.

Long context needs a way to find the right information

Putting more text in a prompt does not ensure that a model will use it well. MIT’s Recursive Language Models keep the input in a code workspace. The model can inspect it, select relevant pieces, delegate smaller questions to further model calls, and assemble the results. Documents become something the agent can work on, rather than something it must absorb in one pass.

The approach lets a fixed model tackle inputs that are too large for a single call. Training on short tasks also transferred to longer inputs and new domains when they shared a useful way of decomposing the problem. The benefit depends on that decomposition, and additional calls can increase cost. Choosing what to read is part of the agent’s work.

Further reading: slide 9.

Skills and memory preserve useful experience

Skills package reusable instructions and code. Memory preserves information for later tasks. Both allow an agent to benefit from prior work without changing the underlying model’s weights. A tested procedure for checking a spreadsheet or a record of a project decision can save the next run from rediscovering the same information.

The production loop in the report explains how to improve these systems deliberately. Record the task, tools, outcome, and corrections. Turn recurring failures into repeatable tests, then decide whether to change context, memory, routing, a tool, or the model. Recording whether each action succeeded gives the team evidence for the next change. Accumulating transcripts without evaluating what worked does not establish learning.

Further reading: slide 10, slide 114.

Better models can need fewer behavioral instructions

Some harness features compensate for weaknesses that a later model no longer has. The report describes Claude Code removing 80% of its system prompt for advanced models without measurable loss on coding evaluations. Examples intended to teach tool use can also constrain exploration when the model already understands the tools.

Harnesses need to be retested as models improve. Their design still needs clear tool interfaces, reliable state, observable actions, and controlled execution environments. After a model upgrade, retest the old planning rules and reminders as well as the model itself. The software surrounding an agent should make work easier to perform and verify.

Further reading: slide 14.

A capable agent can still make poor decisions over months

Long-running work requires an agent to manage resources, judge risk, and learn from previous decisions. KellyBench tests this in a simulated English Premier League betting market. Agents receive a season’s information sequentially and start with £100,000 to build models and size bets. All 12 tested models lost money on average over five runs, and six went bankrupt at least once.

E-CommerceBench gives 18 models CNY 100,000 to operate up to four stores for 365 simulated days. GPT-5.6 Sol produced the highest reported average year-end assets, CNY 1.43 million, while directing 18.48% of order spending to fraudulent suppliers. Opus 4.7 directed 0.12% to fraudulent suppliers. Sixteen of the 18 models showed no clear improvement in bargaining on repeat purchases.

These are simulated markets with particular rules and incentives. They do not measure actual business profitability. They do show that high returns on one measure can coexist with weak supplier screening, and that accumulating experience does not guarantee better decisions. Assessing an autonomous business agent requires examining losses, fraud exposure, and learning as well as its final balance.

Further reading: slide 44, slide 45.

Agent adoption extends beyond developers

What’s exciting is how far this extends beyond developers. OpenAI's study of agent adoption finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.

Non-developer adoption of Codex grows quickly from small starting points.
Non-developer adoption of Codex grows quickly from small starting points.

Models are helping build their successors

AI is already helping build better AI. According to Anthropic's internal index, Claude led 26% of measured model R&D work in August, up from under 1% in February. Researchers set the tasks and supervise execution, and Anthropic reports that this is speeding up development. This is a branch of AI research that I am most excited about.

Claude leads a growing share of Anthropic's measured model R&D, under human supervision.
Claude leads a growing share of Anthropic's measured model R&D, under human supervision.

For many, Karpathy's autoresearch makes part of that loop tangible: an agent edits training code, runs five-minute experiments, and keeps improvements overnight on one GPU. This automates a useful part of research, although sustained, fully autonomous recursive improvement remains to be demonstrated.

What I want to see next is agents developing scientific taste: choosing experiments, recognizing promising directions, and knowing when to abandon a familiar approach. As I argued in Can AI learn scientific taste?, learning that judgment may require the alternatives, failures, and decisions that papers leave out. The ambition is an AlphaGo-like shift in scientific strategy.

Evidence you can use

Same model, different tools and feedback

Study published 2026-05-07

Same model, different tools and feedback
Agent setupTasks resolved
MinimalSource52.5%
ImprovedSource56.5%
FullSource65.5%

GLM-5.1 was held fixed. The table reports mean pass@1 over two runs on a 100-task SWE-bench Verified subset. The increase from the minimal to full setup was 13 percentage points. This is a controlled coding result, not an estimate for every task an agent might attempt.

Sources: Coding study, Table 2.

Frequently asked questions

Answers drawn from the report and the sources below.

What is an agent harness?

The harness is the software around a model that gives it tools, manages context, and handles feedback and failures. Changing these parts can change what the agent accomplishes even when the underlying model stays the same.

Source: Coding study, Table 2.

Can an AI agent reliably run a business for a year?

The report’s simulated markets show substantial remaining weaknesses. All 12 KellyBench models lost money on average, while E-CommerceBench’s highest-earning model also spent heavily with fraudulent suppliers. These tests expose problems in risk management and learning, but do not directly measure real-world business performance.

Source: State of AI Report 2026, slide 44: The house wins: every model loses money on KellyBench sports betting · State of AI Report 2026, slide 45: The highest-earning e-commerce agent is among the worst at avoiding fraud.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. Coding study, Table 22026-05-07. Primary source checked 2026-10-07.
  2. OpenAI: how agents are transforming work2026. Retained from the launch essay.
  3. State of AI Report 2026, slide 7: Same model, better harness = stronger agent2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  4. State of AI Report 2026, slide 8: Agents improve by choosing among specialized harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  5. State of AI Report 2026, slide 9: Recursive language models treat prompts as parts of the environment2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 10: Skills and memory let agents reuse know-how without retraining2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 14: Stronger models can outgrow their harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 44: The house wins: every model loses money on KellyBench sports betting2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 2026, slide 45: The highest-earning e-commerce agent is among the worst at avoiding fraud2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  10. State of AI Report 2026, slide 114: Production feedback guides improvements across the AI stack2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  11. internal index - www.anthropic.comOriginal source link retained from the launch essay.
  12. speeding up - www.anthropic.comOriginal source link retained from the launch essay.
  13. autoresearch - github.comOriginal source link retained from the launch essay.
  14. scientific taste - press.airstreet.comOriginal source link retained from the launch essay.

Cite this page

Benaich, Nathan. “Better tools and context make agents more capable.” State of AI Report 2026. Published 2026-10-08.