I no longer accept AI code until it passes the Gauntlet
Why I ask my coding agent to challenge its work, prove the bugs it finds, and review the fixes before calling a task done.
I am an early adopter of AI assisted coding. Been using it since the days of GPT-4 and when GitHub Copilot was just fancy auto-complete. It’s nice to see how much coding agents have evolved since then - OpenAI’s Codex is my daily driver!
I started building AI agents pretty early - my previous organization had to bring AI capabilities to clients and, thanks to being a part of the Center of Excellence and Innovation team, I got a chance to build their prototypes and see the accepted ones through to production. When you go deep into an agent’s wiring, you will be surprised to see how much of it is regular code. The model does the reasoning, but the harness gives it tools, manages context and lets it interact with its environment. What the harness puts in front of the model, and what the instructions ask it to prioritize, shape the work it does.
Of course, models make mistakes too. But sometimes, the problem is simply that nobody asked the agent to look at its implementation from a different angle. It gets the feature working, runs the tests it thought were relevant and calls it done. Subtle bugs and missed edge cases can survive that entire process.
The review after the review
AI PR reviewers are useful here. Tools like CodeRabbit give us another set of eyes on a change and can point out things the implementing agent missed. I like having that extra review.
But then comes the next part. The reviewer finds a problem, the coding agent fixes it and the tests pass again. Does that mean we are done? The fix might handle the reported example while leaving the same mistake somewhere else. It might even introduce another bug. We need to look at the changed code again.
This got me thinking - why wait until the implementation reaches a PR to start that loop? The agent is already in the repository. It has the task context, can inspect related code and can run reproductions while it is still working.
To be fair, this idea already exists in other tools. CodeRabbit has a CLI that brings reviews into the development workflow too. What I wanted was a review discipline I could give to the coding agent I was already using, through a skill.
Give it something to argue against
In A skill issue, I wrote about how easy it is to lose ownership of code when the agent does most of the writing. One part of that problem is accepting “the tests pass” as enough evidence that the implementation is ready.
Tests cover the cases we chose to write. An agent can make an assumption while implementing a feature and carry that same assumption into its tests. Both agree with each other, and both can be wrong.
Models can come up with plausible objections and alternative scenarios quite readily. I figured, why not give that ability a specific job? Ask the agent to look for a counterexample to the behavior it just implemented.
For example, a CSV exporter works for a table of names and email addresses. Now ask it to preserve zero, false, quotes and different line endings. A retry looks fine when a request fails before reaching the server. Now ask what happens when the server completes the operation but the response never reaches the client.
Those are concrete questions the agent can investigate. “Be more careful” doesn’t give it nearly as much to work with.
The catch is that a model can also argue convincingly about a bug that doesn’t exist. So finding something suspicious is only the start. It has to back the finding up.
Enter Gauntlet
That is how Gauntlet came about. It’s a free, open-source SKILL.md workflow that asks the agent
to review adversarially, fix confirmed problems when authorized and then review the changed implementation again.
The workflow goes through seven steps:
- Map the expected behavior, boundaries and related paths.
- Attack those assumptions with counterexamples and failure scenarios.
- Substantiate findings with a reproduction or a precise argument from the source.
- Trace the underlying cause and check where else it applies.
- Repair that cause when the task allows changes.
- Verify the repair against the changed code.
- Re-attack from a fresh angle.
Basically, the agent has to question its work. If it changes something, it also has to question the fix.
When the host supports sub-agents, Gauntlet can split the review across them. One might inspect data handling while another checks failure paths. They get a defined scope and look for specific ways the implementation can break. The main agent still has to substantiate their reports and resolve conflicting findings. More agents don’t automatically make a conclusion correct.
There is also a solo workflow for hosts without delegation. The skill guides the agent through separate review roles and asks it to be honest about which approach it used.
A small example, with actual runs
I put a controlled CSV exporter example in the repository to show what this looks like. The first draft was deliberately incomplete. The sub-agent reviews, failing checks, repairs and fresh review were actually run, and the evidence is saved alongside it.
That first implementation passed its three baseline assertions. Two sub-agents then reviewed it from different angles: preserving cell values and correctly encoding CSV.
They found two defects. Zero and false became empty cells because the encoder used a truthiness fallback:
const text = String(value || '');
Both values are valid data, but || treated them as missing. The repair used ?? so that only null and undefined took the empty-string
fallback. The other defect was a missing carriage return in the set of characters that required a cell to be quoted.
Both fixes belonged in the shared cell encoder. Special-casing the example row would have left the underlying mistake in place.
The regression checks were run against the saved original too. They failed for those two defects, then all nine passed on the repaired version. After that came a separate review round using new inputs and a separate CSV parser to check that the exported values survived a round trip.
This is the part I care about. The round that made the changes couldn’t also declare itself the clean review. The repaired code had to go through another attack.
It is a small authored demonstration, so I wouldn’t use it to claim how much better Gauntlet makes every coding agent. That needs controlled comparisons across actual agent runs. It does make the intended workflow inspectable.
Keep the findings honest
I don’t want the agent to invent issues just to satisfy an adversarial review prompt. A valid result can be that it found no confirmed defect within the scope it checked.
A finding needs evidence. That can be a failing reproduction, a trace through the relevant code or a bounded argument against a clear contract. Running a test is useful, but the test itself also needs to check the right behavior. If something can’t be verified, the agent should leave it unresolved and explain why.
Gauntlet also asks the agent to track findings, checks and the code revision they apply to. Once the code changes, earlier evidence may need to be collected again. The optional tracker helps keep those records straight; the agent still has to do the actual checking.
There is a finite review budget too. Otherwise, “review recursively” could turn into an expensive loop where the agent keeps restating the same concerns. Further rounds need a reason, such as an unresolved finding or a new way to attack the implementation. If the budget runs out, it should tell me what remains.
And the task still decides what the agent is allowed to do. Asking for a review means review only. Asking it to implement a feature and fix defects allows those changes within the task’s scope. Deployment or unrelated changes need their own authorization. These are instructions for the agent; the host still owns tool permissions and access.
Using it while the agent works
You can install it through the Skills CLI:
npx skills add aakashH242/gauntlet --skill gauntlet
Then, for example, give Codex a task like this:
Finish implementing the CSV exporter. Use $gauntlet to adversarially
review the implementation with sub-agents if available. Reproduce
suspected bugs, fix confirmed defects at their root cause, and review
the changed code again from a fresh angle. Stay within the review
budget and report anything that remains unresolved.
That is the workflow I wanted: implement, challenge, repair, challenge again, then report what was actually checked.
I still have to review the work and own what I ship. But I want the agent to spend some of its effort trying to break its implementation before it hands it back to me. I already use it to write the code. Might as well ask it to argue against the code too 😄
Gauntlet is on GitHub, under the MIT license. If you try it, I’d be interested in the cases where it finds something useful, and just as interested in the ones where it misses a defect or chases a false alarm. Those are the runs that will help improve it.