Better tools and context can help. In a controlled coding study, changing the tools, context, and feedback available to GLM-5.1 lifted success from 52.5% to 65.5% on 100 SWE-bench Verified tasks, without changing its weights. Yes, this is an “older model,” but the principle holds.
Source: Coding study, Table 2.
Improving agents through tools, context, and feedback
As models and the scaffolding around them improve together (and build themselves), I expect many of today's weaknesses to be learned away. We can also get more from capabilities that are already available.
What’s exciting is how far this extends beyond developers. OpenAI's study of agent adoption finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.
Frequently asked questions
Answers drawn from the report and the sources below.
The harness is the software around a model that gives it tools, manages context, and handles feedback and failures. Changing these parts can change what the agent accomplishes even when the underlying model stays the same.
Source: Coding study, Table 2.
The reported gain is 13 percentage points, from 52.5% to 65.5%. Relative to the starting score, that is about 25%. Keeping those two measures separate avoids understating or overstating the result.
Source: Coding study, Table 2.
OpenAI’s adoption study finds the fastest growth among non-developers, including people in legal, sales, recruiting, and marketing. Helping them put these tools to work is a much larger opportunity than the name “coding agent” suggests.
Source: OpenAI: how agents are transforming work.