Guardrails for coding agents · Part 5
I wrote the rule three times. A hook finally held.
Some rules protect things that cannot be undone. Those rules do not belong in prose. They belong in a check the agent cannot talk its way past.
Every part of this series so far has been about rules that shape how work happens: what to check, when, and who decides. This last part is about a different class of rule, the ones that protect something that cannot be undone. For those, I learned the hard way that prose is not enough, no matter how many times it is written.
A rule that was only words
The rule itself is mundane and specific to how our production database is accessed. The details do not matter here. What matters is its shape: a particular command is safe on a laptop and harmful against production, the harm is not visible to the agent that runs it, and the effect lands on real users.
I wrote the rule into the agent's memory in September, after the first time. I wrote it again, in stronger words, into a repository document at the start of October. The agent's own notes record the reasoning. And in the first week of October, a parallel agent working in its own worktree ran the command anyway. The rule was in three places. The agent was in a fourth: a fresh context, briefed for its lane, with no reason to go looking for a rule about a command it considered routine.
That is the pattern worth naming. A written rule reaches an agent only if the agent reads it, and an agent reads what its context contains. The more sessions and subagents you run, the more contexts exist that the rule is not in.
What a hook is
Coding agents call tools: a shell, a file editor, a set of service integrations. The harness that runs the agent can run a small program before each tool call and veto it. That program sees the exact command or arguments about to execute. It does not reason. It matches.
That is where the rule went. A short Python script, registered on the shell tool, on the service integrations, and on file edits. It does three kinds of thing:
- It blocks the command class that caused the harm, on any database, not just production, because the agent cannot always tell which one it is talking to.
- It blocks production changes outright: environment variable edits, production deploys, promotions and rollbacks, firewall changes. Those go in the report for the human.
- It blocks edits to itself. An agent that can rewrite its own guard does not have a guard.
Reads of production stay allowed. That was the owner's second instruction, right after the first: do not take away the agent's ability to look. A guard that blocks too much gets switched off, which is the same failure as a noisy gate from part 1.
The instructions around the hook are as important as the hook. When it blocks something, change the approach or report to the human. Never rephrase, split or obfuscate a command to get past it. That line is there because an agent that is blocked will, helpfully, try another way to do what it was asked.
It blocked the thing that wrote this
The hook is user-level, so it runs in every repository on my machine, including this blog. Yesterday, setting the site up, the agent tried to run a production deploy from the command line. The hook refused. The agent ran a plain deploy instead, reported the refusal, and left the promotion to me.
As it happens, the hosting platform treats a project's first deployment as production regardless. The hook guards the commands it knows about, not the platform's defaults. I mention it because it is the honest shape of every mechanical guard: it closes the class of mistake it was written for, and it is silent about the ones nobody has had yet. Which is exactly why part 2's rule exists: a gate follows an incident, and the incident is written next to it.
Where guardrails belong
Pulling the five parts together, here is how I would place a rule now, by asking two questions. Can the mistake be undone? And can a machine recognise it?
| A machine can recognise it | Only judgment can | |
|---|---|---|
| Undoable | A reporting check: lint, a budget test, a post-push suite | A memory with the reason attached, and a review |
| Irreversible | A hook that blocks, with an escape hatch that demands a reason | The human decides, and the agent presents options |
Most rules land in the top row, and there the cost of the check matters more than its strictness. Part 1's hundred-second hook and part 2's six-thousand-token injection were both in the top row, both too expensive, both removed. The bottom-left corner is the only place where a block with no override is right, and it should be small. On my machine it is one script.
The bottom-right corner is part 4. Nothing mechanical helps there. It is a decision table and a conversation.
The same shape, inside the product
The agent we ship to users went through the same arc in the same weeks, and I wrote that up separately in A coach for the agent. Every behaviour rule we added to its system prompt regressed something else, so we moved rules out of prose and into cheap per-turn checks that nudge only when they fire, with measurement on every nudge.
It is the same lesson from the other side. A rule in a prompt is prose. It is read, weighed against everything else in the context, and sometimes lost. A check at the boundary is not weighed. It runs.
What I would tell someone starting
- A rule exists only if something enforces it. Prose is documentation of the check, not a substitute.
- Price every check per turn. Time and tokens. Remove the ones that cost more than they catch.
- Add a blocking gate after the second incident, or the first irreversible one. Write the incident next to the gate.
- Verify after the push, in the background, waking the agent only on failure.
- Keep memory with the reason. A rule without its why gets argued away by the next context.
- Classify decisions. Product, structure, data, the agent's voice and production are the human's. Show options as mocks.
- Protect the irreversible mechanically. One small script, at the tool boundary, that the agent cannot edit.
That is the series. If you run agents at any volume and have found a different shape, I want to hear it.