← Writing

Guardrails for coding agents · Part 1

I wrote the rules down. The agent broke them anyway.

The first guardrails I put around coding agents were documents. Here is how they failed, what the failures cost, and the shape of the checks that replaced them.

Since the start of September, most of the code in the five repositories I work on has been written by coding agents. Roughly 3,500 commits and 350 merged pull requests went through my account in five weeks. I did not type most of them.

That volume only works with guardrails, and I did not get them right the first time. This series is the story of the routes I went down, in the order I went down them, with what each one cost and what replaced it. Every part stands on its own, but the order matters, because each guardrail exists because the previous one failed in a specific way.

  1. June to August

    Rules in a document

    Coding standards and patterns written down for the agent. A checker that lived on one laptop. A hook that cost a hundred seconds per shell command. This post.

  2. Early September

    A repo whose maintainer is an agent

    Gate classes, budgets for prompt text, and a rule for when a check may block at all. Part 2.

  3. Mid September

    Verify after the push

    Moving tests and type checks out of the agent's way, and what that did to throughput. Part 3.

  4. Early October

    Decisions need a signature

    A week of regressions from choices the agent made on its own, and the rules on the human side of the loop that came out of it. Part 4.

  5. October 6

    When prose fails, write a hook

    A production incident that a written rule did not prevent, and the deterministic check that now does. Part 5.

Route one: write the rules down

The first thing I did, back in June, was what everyone does. I wrote the rules down.

The main repository got a standards document: fifteen numbered rules about simplicity, where types live, how database tables are named, when a helper function is allowed to exist. It got a set of pattern documents, one per area, each with a numbered list of things to do and not do. The agent's instruction file pointed at both.

15
Numbered coding standards
plus six pattern documents
most
Rules broken anyway
in the first weeks
0
Cost of a rule that is only prose
to the agent, per turn

They read like policy. They were not policy. The agent broke them the way a new hire breaks a style guide: not out of defiance, but because a rule it has read once competes with a hundred local cues in the code it is editing. A document says "do not write single-use helpers under five lines". The file it is editing has four of them. The file wins.

The honest lesson from route one is not that agents ignore instructions. It is that a rule with no check behind it has no cost when broken, so nothing in the loop pushes back.

Route two: a checker

So in July the standards got a checker. A script that reads the staged TypeScript files and flags violations. Deterministic rules, such as a table export with the wrong name or a test outside the test directory, exit non-zero. Heuristic rules, such as "this looks over-abstracted", print a warning for a human to adjudicate.

That distinction turned out to matter more than the checker itself. A heuristic rule that blocks a commit gets switched off within a week, because it is wrong often enough that nobody trusts it. A rule that only prints survives, and gets read.

The checker was wired into a pre-commit hook. And then two things went wrong with where it lived.

It lived on one laptop

The script sat in the agent's local configuration directory, which this repository ignores in git by team policy. So the checker existed on exactly one machine. Every other clone committed unchecked, while the rules sat in the standards document reading like policy. The hook's presence read as enforcement. It enforced nothing for anyone else.

Worse, two improvements to the checker had accumulated on that one laptop with no way to share them.

The fix was boring: move the script into a tracked directory, make the pre-commit hook run the tracked copy, and have a missing checker fail the commit outright rather than skip silently. The local copy became a symlink to the tracked one. One implementation, in git, or it is not a rule.

It ran before every shell command

A second check, for lazy TypeScript (as never, non-null assertions, @ts-ignore), was written as a hook on the agent's shell tool. Its docblock claimed it fired on commit and push, which is its own story. What it actually gated on was "is anything staged". With a dirty index, that is always true. So it ran before every shell command the agent issued, including listing a directory.

100 s
Per shell command
measured on a 107-file merge
~25
Shell commands in one session
so about 40 minutes of waiting
0 s
After the fix
fires only on push or PR creation

Nobody noticed for a while, because the agent does not complain about waiting. It just gets slower and burns more of its session on nothing. I found it by measuring a merge that felt wrong.

The check now runs only when the command is a push or a pull request. Everything else goes straight through.

Route three: the full suite on every commit

The pre-commit hook grew. Formatting, the standards gate, a docs coverage gate, and then the whole unit test suite, which had been moved out of continuous integration to save minutes there. A commit took five minutes or more.

For a human committing twice a day that is a tax. For an agent committing every few minutes, as checkpoints, it is a wall. The agent either stopped committing, which meant a crash lost work, or it learned to reach for the bypass flag, which meant the gate enforced nothing.

Two changes came out of this, and both held:

  1. The pre-commit hook is opt-in and off by default. With it off, nothing checks a commit, and the instructions say so plainly, so the agent runs the checker itself before pushing. With it on, the only sanctioned bypass is an environment variable that requires the justification to go in the commit body. The raw --no-verify flag is banned, because it skips with no reason recorded anywhere.
  2. The suite moved to after the push. More on that in part 3, because it changed throughput more than anything else in this series.

Route four: evidence classes

In August we were migrating onto an agent framework and keeping a document of findings to file upstream: things the framework did wrong, or did not document. Before the first filing round I had the agent re-verify every entry against the framework's own bundled docs.

Two of the eight entries were false. One called a behaviour undocumented. It was the documented guard, on a page in the installed package. The other claimed a feature had no opt-out and no docs. Both existed. A third error was mine: I had described an upstream issue as filed by us when we had only commented on it.

The pattern was clean. Entries tagged as measured on a live system held up. Entries that came from reading code were where both failures happened. So the document got a rule that every entry names its evidence class:

ClassMeaning
MeasuredDriven against a real server and observed on the wire: decoded stream chunks, network traces, ledger rows
Evaluated in codeRead directly at a cited file and line in the installed package, not run
InferredReasoned from adjacent evidence without a direct read or measurement. Flagged as weaker

And one sentence that I have reused everywhere since: a claim that something is undocumented is itself a claim requiring evidence.

The same month, the migration playbook got a line that sounds petty and is not: paste the tally line the test run printed, do not type the number you remember. An agent that reports "all 821 tests pass" from memory is reporting a belief. The pasted line is a measurement.

What survived from this period

Looking back from October, four things from these first three months are still in place:

The next part is about a repository that started from these rules on day one, with an agent as its primary maintainer, and what that let us do about budgets and about when a gate is allowed to block at all.