Self-harness: an AI agent that improves its own harness
The interesting part is not that a model can solve tasks. It is that it can improve the way it approaches them.
Self-harness is an arrangement in which an AI model improves the layer that surrounds it: its instructions, its tool definitions and its working rules. The model's weights stay untouched, and only the way it is used changes. In the Self-Harness paper (arXiv, June 2026), which described the pattern, such a loop raised the pass rate of MiniMax M2.5 on held-out tasks from 40.5% to 61.9%.
In short
- A harness is the scaffolding around a model: tools, instructions, memory and workflow.
- The self-harness loop has three steps: find weaknesses, propose a change, validate it with a test.
- The Self-Harness paper (arXiv, June 2026) reports MiniMax M2.5 rising from 40.5% to 61.9% on held-out tasks.
- A parallel system, HarnessX, reports an average improvement of 14.5% across five benchmarks, with the largest gains where the starting point was weakest.
- The base model does not change, so each change is cheap and reversible.
First, what a harness is
A model on its own only generates text. To work as an agent, it needs a layer that turns its responses into actions: a system prompt, tool definitions, memory, error handling, retry limits and a workflow. That layer is the harness, sometimes also called scaffolding.
Ben Dickson of TechTalks calls the harness the model's operating system and points out that it is what separates tools such as Cursor, Aider, Cline and Claude Code from one another. The same model can produce noticeably different results in each of them, because each one hands it the task differently, calls tools differently and behaves differently after an error.
For a business, the consequence is simple. When you buy an "AI agent", what you are mostly buying is the harness, because the model underneath is often exactly the same one your competitor uses. The difference comes from the wrapping, not from the neural network itself.
AI search assistants are built the same way. The search tool, the step that runs the sub-queries a model writes (query fan-out) and the mechanism that ties an answer to live sources (grounding) all live in the layer around the model, not in its weights.
What self-harness means
In the classic set-up, a human improves the harness. An engineer reads the logs of failed runs, changes the prompt, adds a tool, raises the retry limit. Self-harness hands that work to the model itself: it gets access to designated parts of its own harness, proposes small changes to them and keeps only the ones that pass a test.
The scope is fixed in advance. In the Self-Harness paper, the agent could change the system prompt, the start-up instructions, the execution and verification instructions, the procedures for recovering from errors, and the policies that steer a run, including the tool-error limit and the message budget. It could not swap the model or rewrite the whole system from scratch.
This is a close relative of what we have described as loop engineering: designing the loop an agent works in. The difference lies in who tunes that loop.
We are not improving the model. We are improving the way the model works, and that is enough for the same network to start passing tasks it used to fail.
The loop in three steps
- Finding weaknesses. The model reads the traces of its own runs and lists recurring failure patterns: for example, that it gives up after the second tool error, or that it skips checking its own output. The raw material is execution traces, not anyone's opinion.
- Proposing. For each weakness it drafts several variants of the smallest possible fix within the permitted part of the harness. Small, because the smaller the change, the easier it is to tell whether that change is what moved the result.
- Validating. A candidate goes live only after a test confirms that it does not break tasks which were passing before. Rejected proposals simply disappear, and the active harness stays as it was.
The third step matters most, because it is what separates self-harness from tweaking a prompt by feel. Without a gate that measures the effect of a change on the same set of tasks, every modification is a bet. The authors checked results separately on the tasks the loop learned from and on held-out tasks it had never seen.
What the measurements showed
The Self-Harness paper by Hangfan Zhang and colleagues (arXiv, submitted on 8 June 2026) tested the loop on three models and three benchmarks: Terminal-Bench-2.0, SWE-bench Verified and AppWorld. The score improved in every one of the nine model-benchmark pairs, and the authors report relative gains in overall pass rate of up to 132% (Qwen3.5-35B-A3B on AppWorld). The table below shows Terminal-Bench-2.0, a slice of 64 tasks.
| Model | Tasks the loop learned from | Held-out tasks |
|---|---|---|
| MiniMax M2.5 | from 43.0% to 50.0% | from 40.5% to 61.9% |
| Qwen3.5-35B-A3B | from 15.1% to 36.0% | from 23.8% to 38.1% |
| GLM-5 | from 47.7% to 57.0% | from 42.9% to 57.1% |
A second team, the authors of the HarnessX system, measured a similar effect across five benchmarks, among them ALFWorld, GAIA, WebShop and SWE-bench Verified: an average improvement of 14.5%, with a maximum of 44.0%. Their observation is more interesting than the average itself. The biggest jumps came where the starting point was weakest.
Why it works
A large share of an agent's failures comes not from gaps in the model's knowledge, but from how it was asked and by the frame it was made to work in. The model does not know what format to return its result in. It gives up after the first failed tool call. It does not check what it has done. Or it receives a context so long that it loses the instruction somewhere inside it. Every one of these problems is fixed in the harness, not in the weights.
That also explains the point about the weakest starting point. If the configuration was never properly thought through, simply tidying up the instructions and limits gives more than switching to a more expensive model. The reverse holds too: a well-tuned set-up has less left to squeeze out, because the easy mistakes have already been removed.
The other side of the coin
Both papers are preprints on arXiv, so they have not been through peer review, and the figures come from their own authors. Nobody independent has reproduced these measurements yet. Until someone does, read them as a vendor's claim rather than as the result of an audit.
The Self-Harness authors listed the limitations themselves. Their loop makes bounded changes on a fixed set of tasks and is not open-ended self-improvement. The accepted fixes may reflect failure patterns typical of that particular task set rather than a general weakness of the agent. The whole procedure depends on the quality of the checker: if you cannot automatically tell success from failure, you have nothing reliable to base the gate on. For higher-stakes changes, in their view, the condition "it did not get worse" would be too weak.
Then there is cost. Ben Dickson notes that the tuning phase can be computationally expensive and usually needs a strong model in the role of the one rewriting the harness. The savings arrive only later, in day-to-day operation.
There is also a risk the papers do not address directly. An agent allowed to change its own instructions is a new attack surface: if someone fed it doctored execution traces, they could influence what the agent decides its weakness is. We described mechanisms of this kind when writing about agentjacking (in Polish), and the same caution applies here.
What a small business can take from this
You do not need to deploy an automated loop to benefit from this finding. It is enough to adopt its premise: before you pay for a stronger model, fix what surrounds it. In practice, that looks like this:
- Collect failures. Record the specific cases where your assistant or agent falls over, together with exactly what it did along the way. Without that record you do not know what to fix.
- Turn them into a test set. Twenty real tasks from your own business are enough to check, after every change, whether things are better or merely different.
- Change one thing at a time. One fix, one test run, one decision. Three changes at once give you a result you cannot attribute to any of them.
- Set the frame. How many retries after an error, what to do when a tool does not respond, when to ask a human. These are exactly the elements that proved worth fixing in the loop described above.
- Ask your vendor about the harness. Who maintains it, how they measure its effectiveness, and what happens when the model maker releases a new version. We collected questions like these in a checklist for companies choosing an AI implementation partner (in Polish).
One more conclusion is worth keeping, and it comes from a completely different direction. Research by Anthropic showed that when working with agents, knowledge of your own field beats programming experience (in Polish). A harness is largely a written description of your process, and you know that process better than any vendor does.
Common questions
What is a harness in AI?
It is the layer around the model: tool access, instructions, memory and workflow, which together turn a bare model into a working agent.
How is self-harness different?
The model improves its own harness instead of a human doing it. The base model does not change; what changes is the way it is used.
How does the self-harness loop work?
In three steps: finding weaknesses in the current harness, proposing a specific change, and validating whether that change actually improves the result.
What results does it deliver?
In the Self-Harness paper, every one of the nine model-benchmark pairs tested improved, and in the HarnessX system the average gain was 14.5% across five benchmarks. These are lab results from preprints with no independent replication, so treat them as an upper bound, not a promise.
Sources
Read next
Find out whether AI recommends your company.
Start with the free SEO and GEO audit, delivered in 5 working days. We check how the models describe your brand and hand back a prioritised list of changes.