Measuring my agent harness against plain Claude Code
I built a multi-agent harness and ran it head to head against a plain control on a real benchmark. On these tasks it cost about five times more and did not score better. What it buys instead is verification, reproducibility, and an audit trail, and none of that is what a pass-or-fail grader measures. Here is why that is the answer I expected, and why the benchmark cannot settle the question.
Most of what I have written here argues for structure around a coding agent: separate roles, a verifier that did not write the code, working rules the agent has to follow. I run a harness like that on real work every day. I had never measured whether it beats the thing it wraps.
So I measured it. The result is that on this benchmark the harness cost about five times more and resolved one task fewer than a plain control. I think that is the correct outcome, and I think the benchmark I used cannot answer the question I actually care about. Both of those are worth explaining.
Harness and control
The harness is a main agent on Opus with two sub-agents. A measurer on Sonnet runs the commands it is given and reports the output verbatim. A reviewer on Fable reads the finished diff and returns PASS or BOUNCE, and diagnoses a fix that failed. A written set of rules binds the main agent: read the issue twice, reproduce the bug on the unfixed code before believing any fix, go up the source hierarchy rather than sideways (issue text and the checkout first, memory last), fix at the root, and hand the diff to someone who did not write it before calling it done.
The control is the same model and the same task prompt with none of that. Plain Claude Code, one agent, no sub-agents, no rules file. The only difference between the two configurations is the harness.
Claim under test
The claim I wanted to check is the one the harness exists to justify: that reproduce-measure-review resolves more issues than a single agent given the same model and the same prompt. Not that it is safer or leaves a better audit trail, both of which I already believe. That it scores higher.
Clean-room setup
I used SWE-bench-Live Lite, 300 real GitHub issues from Python projects, each with a known-good fix and a grader. I ran a 20-task pilot, both configurations on every task.
Each task runs in its own container with three things locked down. The checkout contains only history up to the issue, with every later commit scrubbed, so the fix cannot leak in from the project's own future. There is no network except the model API, reached through a proxy that refuses GitHub, PyPI and the git remotes. The grader is the official one, which resolves the known-good patch and rejects a no-op. Before the run I fed the whole pipeline a known-bad case and watched it fail, because a clean result from an instrument nobody proved is not evidence.
Pilot results
Two of the twenty tasks have gold patches that fail on my host, so they score for neither configuration and I set them aside. That leaves eighteen tasks with a valid answer.
| Configuration | Resolved | Cost (USD-eq) | Median minutes |
|---|---|---|---|
| Harness | 9 of 18 | 18.78 | 3.2 |
| Control | 10 of 18 | 3.87 | 0.6 |
The outcomes are the same on nineteen of the twenty tasks. Per rollout the harness cost 0.94 and the control 0.19, so the harness is about five times the spend and about five times the wall-clock time for a result that is, on this set, one task worse. There is no task the harness resolved that the control missed.
That last sentence is the honest ceiling. I cannot tell you the harness found something the plain agent could not. On Lite, it did not.
Memory versus the specification
The single task that separated them is worth reading closely, because it is the whole argument in one case.
On conan-io/conan-17366 the control passed and the harness failed. The control reproduced the project's actual upstream fix from memory and applied it. The harness followed the issue text, which described the desired behaviour but not the same change the maintainers eventually made, and produced a fix the grader rejected.
My rules told the harness to do that. "Never your memory first" is written into them, because on the work I am paid for the model's memory of a library is often wrong and always unverifiable. On this task, memory happened to be right and the specification happened to be incomplete, so the discipline cost me the point. I would keep the rule.
Dataset age and recall
That case is not a fluke of one task. It is the shape of the whole dataset.
Lite's issues are from October 2024 to March 2025. They are older than the model under test, so the model has seen many of these projects at or past the point the issue was fixed. Recall is part of every score, for both configurations. On deepset-ai/haystack-8489 the fix the agent wrote shared the first eight hexadecimal characters of its release-note filename with the real upstream one, which is the model reconstructing a specific commit it has partly memorised. The isolation stops the agent from looking the answer up. It cannot stop the agent from already knowing it.
A benchmark where recall is a free variable rewards the configuration that leans on recall. That is the control. The harness spends its extra turns proving a fix against the code in front of it, which is exactly the work you do not need when you can remember the answer.
Cost and value of the harness
The five-times cost buys three things: a reproduction proven to fail on the unfixed code, a measurement run by a sub-agent instead of asserted by the one that wrote the fix, and a review by a model that did not write it. On a task that a single agent closes correctly in four turns and forty seconds, all three are overhead. The task is small, the specification is enough, and the answer is gradable by a script. Nothing in that description is true of the work I built the harness for.
What those three steps produce is not a better answer. It is an answer that arrives with the reproduction, the measurement and the review attached. The plain agent gives you the fix. The harness gives you the fix together with the evidence it works and a record another person can read, re-run, and get the same result from. On the benchmark those properties score nothing, because the grader already holds the right answer and confirms it in a second. Consistency, reproducibility and an audit trail are worth zero when a script can check the result for free.
The work I built it for is long, spans many files, has an ambiguous specification, and has no grader at the end except a person. That is where an unproven fix and an unverified "done" are the actual failure mode, the one I have written about before. There, the evidence and the record are not overhead. They are what catches a confident wrong answer before it ships, and what lets a second contributor trust a result they did not produce. A benchmark of small, well-specified, machine-graded fixes is the one setting where the harness can only lose. It prices the premium and gives no credit for what the premium covers.
MultiLang 2026 as the real test
The experiment that would mean something is the same harness and control on tasks the model cannot have memorised. SWE-bench-Live has a MultiLang split with 663 tasks from 2026, after the model's cutoff. There, recall is not a free variable, and the question becomes whether reproduce-measure-review beats a single agent when neither can fall back on remembering the fix.
I have not run it yet. It needs the runner generalised past the Python-only path, and it needs a spend cap I am willing to set, because 663 tasks at the harness's rate is real money. When I run it, the number will be worth more than this one, whichever way it goes.
Reading the number
The pilot did not show that my harness is worse. It showed that it is worth nothing on tasks a plain agent already solves cheaply, which is what a five-times premium for verification should do when there is nothing to verify. The measurement I can stand behind today is a null result on the wrong dataset, reported as a null result on the wrong dataset. The one that tests the actual claim is still to run, and I will publish it the same way when it does.