Skip to main content

Command Palette

Search for a command to run...

Combatting Slop with Quality Gates

Updated
•13 min read•View as Markdown
Combatting Slop with Quality Gates

Solving the Hardest Problems Before Writing Code

Three weeks ago, I installed Claude Code on my new MacBook and configured a custom toolset tailored for spec-based development on my side projects.

Instead of jumping straight into writing code, these custom skills manage my entire pre-implementation pipeline:

  • /research: Conducts deep-dive investigations and returns a strictly cited Markdown note.

  • /spec: Translates that research note into an actionable implementation spec, deeply rooted in my existing codebase architecture.

  • The "cold Review": The tool passes the completed spec to a fresh subagent with zero knowledge of the prior discussion, explicitly instructed to pull the plan apart.

If the cold review subagent flags an ambiguity it cannot resolve by reading, the issue escalates to a Spike - a time-boxed, disposable experiment run on a temporary git branch. To stay within token limits and maintain pristine focus, each skill automatically spawns a dedicated subagent whenever it requires a fresh context window.

These skills can be found on the Claude Code marketplace and in this GitHub repo. This diagram depicts the full workflow that this set of skills enables:

Where the skills could be stronger is in the area of preventing the creation of slop code.

How do we tell when code is slop?

Quite a few people complain about slop, but fewer people can actually nail down what slop is in a quantifiable manner. A quantifiable definition is all-important, because it determines what can catch it.

Slop is plausible code that should not exist. A wrapper around a single function call. An abstraction for a second case that never arrives. A helper duplicating one three lines above it. A comment restating the line below it. An exception caught and silently dropped. None of it is wrong, exactly. It passes review if you skim. It simply should not be there, and six months later no one can identify which parts were load-bearing.

This reminded me of the software engineering track on my Computing Science degree course, specifically the part on code quality metrics. My first thought was that metrics provide the answer. It turns out that they provide part of the answer, but not the answer in its entirety.

Pillar one: metrics for code quality

Classical code quality metrics include the likes of:

  • Cyclomatic complexity measures the number of linearly independent paths through code - Thomas J. McCabe's 1976 metric, and still central to most quality tools.

  • Test coverage reflects the portion of code that tests actually run.

But neither captures slop. A compact hand-written parser scores badly on cyclomatic complexity yet remains excellent code. Forty lines of unnecessary indirection score well and are pure slop. Coverage is worse: code generated alongside its own tests shows high coverage, because the agent produced both.

This is why I am drawn to the tools that Robert C. Martin - "Uncle Bob", of Clean Code fame - has developed around SwarmForge. The notable aspect of SwarmForge is not its metrics as such but the fact that they are expressed as roles. It operates agents in tmux sessions, each in a separate git worktree assigned a task from a craftsmanship playbook: a specifier writing Gherkin, a coder using TDD, a cleaner performing DRY and CRAP reviews, an architect protecting dependency direction, a hardener running mutation testing, and QA checking the results.

The metrics it relies on are:

  • CRAP score - Change Risk Anti-Patterns, created by Alberto Savoia. It combines a function's cyclomatic complexity with its test coverage, so a function that is both tangled and poorly tested scores badly. It stands out among the classical metrics because it uses a conjunction. Martin has built language-specific tools for it (crap4java, crap4clj, crap4go) and more recently crapper, which handles Clojure, Java, Go, TypeScript, Rust and Python together.

  • Duplication, through his dry4* tools.

  • Coverage, required as an input for CRAP.

  • Module size, estimated from the mutation tool's possible changes - a sign that a file may handle more than one task and could be split.

Uncle Bob's own experiment argues against a strict complexity cap - so the person whose tooling I am adopting warns against hard-capping one of CRAP's two inputs. I think he is right, and that CRAP survives the objection where raw complexity does not: a conjunction of two signals is much harder to game than either alone.

I am also drawn to SlopCodeBench, a community benchmark that measures "code erosion as agents iteratively extend their own solutions across checkpoints". Its landing page puts it this way:

Beyond correctness, we measure code erosion — verbosity, dead branches, and redundant structure — to surface the agents that stay clean under sustained change instead of patching their way into slop.

In summary, metrics tell part of the tale, but they do not see the duplicated helper or the comment restating the code. Something else has to.

Pillar two: linters catch what metrics can't

A linter is a static analysis tool that identifies code smells, probable bugs and stylistic issues without executing code. It is unglamorous, and it is the actual workhorse of slop prevention, because metrics measure the shape of code while slop is mostly a matter of pattern.

Consider three ruff rules I am adding to CI:

  • S110 - try/except/pass. The exception that is quietly ignored.

  • BLE001 - a blind except. Catching everything turns a genuine bug into a puzzle.

  • ERA001 - commented-out code. Agents often leave this behind.

None of these alter a complexity score. All three represent slop, all are deterministic, and all block a build without requiring human judgment. That final point makes them valuable in an agentic loop: the agent receives a red check and a rule name rather than a vague impression.

Pillar three: the harness

The third pillar concerns how you configure the tool performing the work. It begins with CLAUDE.md (or AGENTS.md for a vendor-neutral option). "Karpathy's CLAUDE.md" is a good starting point - a community summary of Andrej Karpathy's public comments on how LLMs fail at coding, condensed into a four-rule file. Refer to the repository swarmclawai/andrej-karpathy-skills for the file. The rules are:

  • think before coding

  • simplicity first

  • surgical changes

  • goal-driven execution

The file also includes a six-rule self-check extension attributed to Karpathy.

I applied my /research skill to the vendor guidance and found convergence on four points:

  • provide the agent with a runnable check instead of a principle

  • keep rules brief

  • direct it toward an existing pattern in the codebase rather than describing the pattern in prose

  • enforce via tools rather than wording

Uncle Bob's conclusion is more direct: "You can't tell an agent to be clean."

That is the whole case for gates over instructions. It is worth saying that my own research rated itself honestly - medium confidence on the ranking, low on effect sizes, and no evaluation grading generated code yet.

A worked example: the check that found what the checker couldn't

Here is a worked example that convinced me a cold review earns its keep, irrespective of what I decide to do going forward.

The /spec skill runs a verifier over the completed spec: it checks each path:line citation against the repository, each number against its source, and whether each work item's acceptance criteria can be observed. It ran three times. The final round confirmed 54 of 55 claims and judged the plan sound.

Then the cold review, the adversarial agent that had seen none of the conversation, identified two issues the verifier could not detect:

  1. A work item's acceptance test loaded two documents and deleted one. Removing one of two is 50%, and the pipeline has a safety guard that rejects deletions above 20%. The test could never have passed. The spec's own arithmetic made its acceptance criterion impossible.

  2. The spec's documented recovery path for a withheld deletion failed entirely. I verified this by running it: three full load–evaluate–promote cycles, and the table still held all ten documents it was meant to shrink to five.

I corrected both, recorded them, and ran another adversarial pass over just the changes. It detected the same class of error again - another acceptance test rendered impossible by the same 20% guard.

Twice. Same guard. In a spec that a deterministic checker had passed with 54 of 55.

The verifier acted as a deterministic gate and performed its role well - every citation was confirmed. It simply did not test whether a number made a test unpassable. A different check did. And that is the point underneath all three pillars: different kinds of check catch different classes of problem, a passing checker is not the same as a correct result, and no prompt instruction would have produced either finding.

Where to next for my skills

My initial thoughts centred on some form of gate that would cause the /implement skill to loop until the work it was carrying out on a fresh branch met the pass threshold.

Here, paraphrased, is what the research note recommended. The order is by how much evidence there is that the change works, not by how much it would help - which is why widening the linter sits at five despite being, by the argument above, the most useful of the lot.

  1. Check the clean-up step too. After the implementer writes code, /simplify and /code-review --fix tidy it - and those passes can stray. Re-run the scan afterwards and revert any clean-up commit touching a file outside the work item's declared list. Anthropic has documented the code-review skill, at maximum effort, making out-of-scope edits, so this is the item with the hardest evidence behind it.

  2. Scope the implementer's instructions. Add explicit wording that it should do only what the work item asks, and stop and report rather than widen the job.

  3. Catch more kinds of slop automatically. Beyond my draft /implement gate already scans for: swallowed exceptions, an unlogged new third-party dependency, and comments that only restate the code.

  4. Give each work item a pattern to follow - a path:line pointing at existing code that already does the thing properly, rather than describing it in prose. The spec verifier already checks citations, so a false pointer gets caught.

  5. Widen the CI linter - S110, BLE001 and ERA001, once I have counted the existing hits.

  6. Three checkable lines in CLAUDE.md - prefer an existing helper, never catch-and-ignore, no comments restating code. Last on purpose: it is the weakest lever, and the one Uncle Bob's objection lands squarely on.

None of that contradicts my initial instinct so much as sharpens it. The loop is the right shape; what the research adds is the content of the gate inside it - items three and five say what it should actually measure - plus a step I had not thought of at all, which is that the clean-up pass needs watching as closely as the implementation.

Takeaways

Whatever I end up building, these are the things I will carry into it:

  • Metrics catch risk, not slop. Complexity tells you where change is dangerous. It says nothing about the helper duplicating the one three lines above it, or the comment restating the line below. Linters do more of that work than the metrics I started out convinced by.

  • Coverage is the worst metric for agent-written code. The agent wrote the tests as well, so high coverage on generated code tells you almost nothing.

  • Only the rules a tool enforces survive. Everything else is a polite request, and Uncle Bob is right that you cannot ask an agent to be clean.

  • A passing deterministic checker is not a correct result. Mine confirmed 54 of 55 claims and still let through two acceptance tests that could never have gone green.

  • Use more than one kind of check, because they fail in different directions. A deterministic verifier and an adversarial reader caught completely different classes of problem, and neither substitutes for the other.

Who to follow in this space

Researching this post turned up a lot of people worth following. Here they are in one place, so you do not have to assemble the list yourself:

  • Birgitta Böckeler - Distinguished Engineer at Thoughtworks and author of the Exploring Generative AI series on Martin Fowler's site. Her work is the closest thing to prior art for this post: one memo covers maintainability sensors for coding agents, another covers context engineering, with Claude Code as the worked example. Practitioner reports, not predictions. martinfowler.com/articles/exploring-gen-ai.html

  • Kent Beck - popularised TDD, and now writes about "augmented coding", which he separates from vibe coding on precisely the grounds this post argues: you still care about quality, complexity, coverage and design. Start with Augmented Coding: Beyond the Vibes. newsletter.kentbeck.com

  • Robert C. Martin - the craftsmanship end of the argument, and the source of the tooling I have been quoting. Read the repos as well as the posts: he ships the actual CRAP, DRY and mutation tools rather than opinions about them. github.com/unclebob · @unclebobmartin

  • Simon Willison - the most reliable running log of what these tools do, rather than what their vendors say they do. He tries things, publishes the transcript, and is unusually honest when something fails. simonwillison.net

  • Andrej Karpathy - infrequent, high-signal, and the source of the observations those four rules were distilled from. Follow him for the failure modes he names before anyone else has a word for them. @karpathy

  • Armin Ronacher - the creator of Flask, writing about agentic coding from inside real projects. Good on the parts that are genuinely harder than the demos suggest. lucumr.pocoo.org

  • Thorsten Ball - builds agents for a living and explains the mechanics plainly. Read him if you would rather understand your harness than treat it as a black box. registerspill.thorstenball.com

  • Matt Pocock - the spec engineering angle, and the clearest short-form explanations of agent tooling going. aihero.dev · @mattpocockuk

A caveat on shelf life

This space is moving at a furious pace, and some of the content in this post will date - and quickly. Rapid model improvements since the turn of the year have already led "Uncle Bob" to relax his strict agent testing gauntlet. Meanwhile, industry luminaries continue to drop fresh insights almost daily:

Closing thoughts

I agree with Uncle Bob's "Bathrobe rant on slop": AI coding harnesses can certainly act as slop cannons, but that outcome, with a little effort, is far from inevitable. The bad press reflects adoption moving faster than the engineering discipline, and that discipline largely relies on familiar tools - metrics, linters, gates, adversarial review - applied to a new kind of author.

You may have encountered the terms "prompt engineering" and "harness engineering" - add "verification engineering" to your toolbox of AI engineering competencies. This is the skill of constructing the checks that confirm whether a harness delivered what you requested. Despite much doom-mongering about coding harnesses killing off software engineering jobs, engineering as a discipline will survive, but its centre of gravity will shift.

Ultimately, this post comes down to two main conclusions:

  • You cannot instruct an agent to be clean. You can only decline to merge what fails the test.

  • As models evolve, so should your engineering discipline.

A

A gate is only as good as what the spec says when it's wrong. That part stayed open. If the spec is the thing slop gets measured against, what happens when the agent satisfies every check and the feature is still the wrong one? Can the gate fail a spec, or only the code written against it?

C

I thought I had already replied to this, but thanks for the comment by the way. The short answer, in my humble opinion, is that there will be some gates that may be tightly coupled to a spec, and some that are spec-agnostic - a gate associated with a linter, for example. I'm currently looking at a publicly available repo, getting my harness (Claude Code) to recreate each PR based on a spec derived from each PR's documentation, and then measuring both the original PR and the harness-generated one against CRAP tests, mutation tests, linter output, the metrics used by SlopCodeBench, etc. The art in doing this is to do it in such a way that the harness does not realise what the test is and game it by copying the original PRs verbatim.