One task
You describe the change and the proof you expect.
Smart Harness catches candidates that look finished but fail your checks, gives the model a bounded chance to recover, and applies the verified winner. You see every attempt, every token, and the final verdict.
candidate generated
readyverifier: 1 test failed
retrybounded recovery attempt
metered7 / 7 tests passed
verifiedThe public contract is simple. Your checks—not the model's confidence—decide when the task is done.
You describe the change and the proof you expect.
The selected policy controls how far recovery can go.
Tests or checks decide whether a candidate is acceptable.
A verified winner is applied once—or the result says verification failed.
No secret sauce diagram. Just the customer-visible contract: bounded work, independent checks, and an honest result.
Step through a realistic boundary-bug repair. The left side keeps going when the verifier finds evidence; the right side hands that loop back to you.
Same repository, same model, same failing test, and the same definition of done.
The single-pass agent starts from the identical prompt and repository state.
On a controlled 1,140-task BigCodeBench run using production infrastructure—not organic customer traffic.
627 of 1140 passed
817 of 1140 passed
190 more verified completions
billing-invariant violations
Controlled benchmark result. It does not predict the lift on every repository, model, policy, or task.
These are direct captures from a real Verdict Desktop code-task completion. They show the finished workspace and final response receipt—not a Smart Harness A/B screenshot.


A failed first candidate can trigger more calls. Smart Harness makes that trade visible instead of hiding it behind a success animation.
Your selected policy caps the number and kind of additional attempts.
Tokens, wall time, and the relevant cost or energy signal roll into the outcome.
The proven benefit is more verified completions and less manual retry choreography—not universal token or time savings.
No. It is a verification and recovery layer around the model and task you choose.
No. Every policy has a bounded attempt budget. You can choose a faster or more thorough policy, and the outcome reports what actually ran.
No. A failed candidate can trigger extra attempts. The gain is a higher rate of verified completions and less manual retry choreography; every attempt is metered.
The repository-aware engine runs in Verdict Code and in Verdict Desktop portable coding tasks. Other surfaces may expose catalog controls without running the full code-task engine.
Start locally for $0. Turn on Smart Harness when the task deserves a verified result.