We'd built a tool to keep our AI agents from creating piles of low-quality code, but when I found a bug in kragg, I started doubting the code it had already let through.

A lot of people are claiming 10x or even 100x productivity gains from coding agents. But if engineers are producing 10x as many changes, nobody is going to suddenly start reading 10x as much code. If you have a small team, the more changes you produce per day, the less scrutiny each change gets. Eventually code gets shipped without review because everyone assumes somebody, or some agent, looked at it.

So we built a tool betting on that trend.

kragg is designed around the idea that code may get checked in without a human carefully reading it line-by-line first. Instead of asking another agent to stare at the diff and invent concerns, we define rules that our tool enforces. kragg runs those checks, sends failures back to the coding agent, and keeps doing that until the code meets the standard.

We deliberately don't use agentic code review as the primary gate. In our experience, if you ask a model to find problems, it will find problems. ESPECIALLY, if you use adversarial agentic review. On top of that, the results are slower, harder to reproduce, and more often than not, very subjective. Kragg's job is much smaller: make the acceptance criteria explicit, make the checks repeatable, and let the agents fix their own work before it lands.

Unfortunately, the first versions of our tool had a paperwork problem.

kragg could find an old report saying the work had passed, and use that to approve the new work. It’s like going to the DMV to renew your registration with a 3yo inspection report and the clerk accepting it as current.

The change to require a fresh result was trivial. I recreated the failure with the old report still sitting there and confirmed it no longer earned a pass.

But there was this lingering question bothering me: what about the work we'd already approved and shipped?

This is where having kragg built into our agents' everyday work paid off. As they worked on the code again, they ran the updated checks. When those checks found problems, the agents corrected them. I didn't have to go through everything by hand and turn each defect into another assignment.

If your team keeps giving AI the same feedback over and over, don't add another reviewer. Turn that feedback into a check that runs every time.

If you're using Python, download Kragg on GitHub, wire it into your coding agents, and start with one rule you never want to enforce by hand again.

And if your team is spending more and more time fixing AI's work, but you're not sure what can actually be automated, talk to us. We can look at your workflow, figure out what can be caught earlier, what still needs a person, and where Torta can actually help.

— Alan, Colin & the Torta team

P.S. For a broader look at running agents in production, read our AI Ops field guide. It covers approvals, failures, and questions to ask about your own setup.

Reply

Avatar

or to participate