Writing · 2 min read

Your AI agent should not decide when it is done

A local coding agent with access to tests still reported broken code as fixed in 8 of 40 tasks. One harness rule, finished means tests pass, brought that to zero.

The standard pitch for AI agents is that they check their own work. Give a coding agent the ability to run tests, and it will keep going until the tests pass. I measured whether that is true, with a local model (Qwen3-8B, on one consumer GPU) fixing 40 small buggy programs from the public QuixBugs benchmark, in a locked-down sandbox. The full setup is on the LAB/003 page.

Running the tests barely helped

Programs fixed, out of 40Graded by the same tests, offline, after the agent stops

Qwen3-8B Q4_K_M

  1. A · one-shot, no tools62.5%
  2. B · agent, cannot run tests60.0%
  3. C · agent with test feedback67.5%
  4. D · C, and only passing tests end the task70.0%

Qwen3-14B Q4_K_M

  1. A · one-shot, no tools75.0%
  2. B · agent, cannot run tests52.5%
  3. C · agent with test feedback67.5%
  4. D · C, and only passing tests end the task82.5%

The agent that could run tests (C) fixed 67.5% of the programs. The same model answering once, with no tools at all (A), fixed 62.5%. Two tasks out of 40 is within noise.

Because the agent stopped on failing tests

The traces explain it. In all 8 tasks the agent got wrong in condition C, it declared the work finished without a passing test run. Six times it made a last change and stopped without testing it. Twice it stopped straight after a failing run, describing a fix it never made. It told me twice that the original code “was already correct”, and once that the tests were probably wrong.

None of this is unusual model behaviour. A language model produces a plausible end to the conversation. “The fix is complete” is a plausible end. Whether it is true is a different question, and the model is not the right component to answer it.

One rule fixed it

In condition D, the harness refused “done” unless the most recent test run, after the last change, had passed. Otherwise it told the agent the tests had not passed, within the same budget of 8 turns.

  • False “fixed” results went from 8 to 0. Every unsolved task ended as a clearly reported failure.
  • With the larger Qwen3-14B, the same rule lifted the success rate from 67.5% to 82.5%, the best result in the Lab.

The rule is a few lines of code. It does not make the model smarter. It changes who decides.

This is not only about code

Any agent that acts needs a definition of done that lives outside the model:

  • Code: the tests pass, the build succeeds, the linter is clean.
  • Documents: the extracted totals reconcile; the required fields validate (the same idea as the arithmetic checks in LAB/002).
  • Operations: the record exists in the target system with the expected values; the API returned success, and a read confirms it.

When that check cannot be written, that is the signal that a person must review every result, and the system should be designed and priced that way.

What to measure

The headline number for an agent is usually its success rate. The number that decides whether you can trust it is the false-success rate: how often it says “done” when it isn’t. Measure that one first, and design the harness so it is zero, because a reported failure costs minutes while a false success gets shipped.