All topicsSTATE OF AI REPORT.

Can an AI agent improve without changing its model?

Better tools and context can help. In a controlled coding study, changing the tools, context, and feedback available to GLM-5.1 lifted success from 52.5% to 65.5% on 100 SWE-bench Verified tasks, without changing its weights. Yes, this is an “older model,” but the principle holds.

Source: Coding study, Table 2.

Evidence you can use

Same model, different tools and feedback

Study published 2026-05-07

Same model, different tools and feedback
Agent setupTasks resolved
MinimalSource52.5%
ImprovedSource56.5%
FullSource65.5%

GLM-5.1 was held fixed. The table reports mean pass@1 over two runs on a 100-task SWE-bench Verified subset. The increase from the minimal to full setup was 13 percentage points. This is a controlled coding result, not an estimate for every task an agent might attempt.

Sources: Coding study, Table 2.

Improving agents through tools, context, and feedback

As models and the scaffolding around them improve together (and build themselves), I expect many of today's weaknesses to be learned away. We can also get more from capabilities that are already available.

What’s exciting is how far this extends beyond developers. OpenAI's study of agent adoption finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.

Frequently asked questions

Answers drawn from the report and the sources below.

What is an agent harness?

The harness is the software around a model that gives it tools, manages context, and handles feedback and failures. Changing these parts can change what the agent accomplishes even when the underlying model stays the same.

Source: Coding study, Table 2.

Sources and dates

2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.

  1. Coding study, Table 22026-05-07. Primary source checked 2026-10-07.
  2. OpenAI: how agents are transforming work2026. Retained from the launch essay.
  3. State of AI Report 2026, slide 7: Same model, better harness = stronger agent2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  4. State of AI Report 2026, slide 8: Agents improve by choosing among specialized harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  5. State of AI Report 2026, slide 9: Recursive language models treat prompts as parts of the environment2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  6. State of AI Report 2026, slide 10: Skills and memory let agents reuse know-how without retraining2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  7. State of AI Report 2026, slide 14: Stronger models can outgrow their harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  8. State of AI Report 2026, slide 44: The house wins: every model loses money on KellyBench sports betting2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  9. State of AI Report 2026, slide 45: The highest-earning e-commerce agent is among the worst at avoiding fraud2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
  10. State of AI Report 2026, slide 114: Production feedback guides improvements across the AI stack2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.

Cite this page

Benaich, Nathan. “Can an AI agent improve without changing its model?.” State of AI Report 2026. Published 2026-10-08.