Same model, different tools and feedback
Selected evidence from the State of AI Report 2026, with the period, source, and definitions needed to interpret it.
Evidence you can use
Same model, different tools and feedback
Study published 2026-05-07
| Agent setup | Tasks resolved |
|---|---|
| MinimalSource | 52.5% |
| ImprovedSource | 56.5% |
| FullSource | 65.5% |
GLM-5.1 was held fixed. The table reports mean pass@1 over two runs on a 100-task SWE-bench Verified subset. The increase from the minimal to full setup was 13 percentage points. This is a controlled coding result, not an estimate for every task an agent might attempt.
Sources: Coding study, Table 2.
Sources and dates
2026 report snapshot. Preview revised 2026-10-07. Individual data periods and source checks are listed below. This is not a claim that every source was updated on that date.
- Coding study, Table 22026-05-07. Primary source checked 2026-10-07.
- OpenAI: how agents are transforming work2026. Retained from the launch essay.
- State of AI Report 2026, slide 7: Same model, better harness = stronger agent2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 8: Agents improve by choosing among specialized harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 9: Recursive language models treat prompts as parts of the environment2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 10: Skills and memory let agents reuse know-how without retraining2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 14: Stronger models can outgrow their harnesses2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 44: The house wins: every model loses money on KellyBench sports betting2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 45: The highest-earning e-commerce agent is among the worst at avoiding fraud2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
- State of AI Report 2026, slide 114: Production feedback guides improvements across the AI stack2026 report snapshot. Read against the report PDF on 2026-10-07. Study-specific limits retained.
Cite this page
Benaich, Nathan. “Same model, different tools and feedback.” State of AI Report 2026. Published 2026-10-08.