Guardrails for coding agents · Part 2
When the agent is the maintainer
Starting a codebase where agents are the primary maintainers from day one. Gate classes, budgets for prompt text, and a rule for when a check is allowed to block.
In early September we started a new monorepo from scratch. The first line of its agent instructions says who it is for: the primary maintainer of this repository is AI coding agents, and every convention is written for them.
That sentence changed what the guardrails could be. In an existing codebase, every check has to tolerate years of code that predates it. In a fresh one, a check can be absolute from the first commit. Part 1 was about rules that failed because nothing enforced them. This part is about what enforcement looks like when you get to design it up front.
Three classes of gate failure
The first document we wrote was not a style guide. It was a one-page policy on which checks block what, and when. Its opening line is the whole idea: a gate that stops a from-scratch build from moving is a cost, not a safety; a gate that lets a bad row reach a shared database is the opposite.
Every gate failure falls into one of three classes, and the class decides the response.
| Class | Examples | What happens |
|---|---|---|
| Auto-fixable | Formatting, import order, a trailing newline | The hook fixes it and re-stages. Never a red light. |
| Cheap to fix later | A doc over its size budget, a hard-wrapped line in a prompt file | Reported in the run and as a tracker row. Fixed in the next docs commit. |
| Fix now | A secret in the diff, a compile error in the agent's own files, a broken link or a cited path that does not exist, a failing test in a package the commit touched | The commit stops. |
The test for "fix now" is whether the failure misleads the next agent or cannot be undone. A wrong citation sends someone to a file that is not there. A leaked secret is out. Everything else is a note.
Three things block at every stage, no exceptions: a tool input schema that carries an account identifier, a write to a remote database, and a push to the main branch.
When a new gate is allowed
The policy also says when a gate may be added, and I think this is the most useful sentence in the whole series:
Add a gate when a class of mistake has happened twice and a check can catch it in under a few seconds, or when the mistake is irreversible. A gate with no incident behind it is a guess and starts as a report, never a block.
Every incident that produces a gate gets written into a lessons file with the gate that now prevents it. The gate carries its own justification. When someone asks "why does this block me", the answer is a date and a mistake, not a preference.
A pre-commit hook that finishes in seconds
The pre-commit hook runs only what finishes in under twenty seconds and would embarrass a reviewer if missed: format the staged files, type-check the packages whose files are staged, compile the agent definition when one of its files is staged, and block outright on a secret-shaped string in the diff.
Tests, evals and the slower repository gates run after the push, which is part 3.
The hook has an escape hatch, and the escape hatch has a rule. A skip is an environment variable that requires a non-empty reason, and the reason is printed. Skipping with an empty reason is refused. The raw --no-verify flag is not used, because it skips with no reason recorded anywhere, and a skip you cannot trace back is the one that bites you.
Budgets for prompt text
An agent framework turns files into a prompt. Instruction files load every turn. Tool descriptions load every turn the tool is mounted. Skill descriptions load every turn so the model knows what it could load. All of that is paid for on every single call.
So the repository has budgets, enforced by a test that fails when they are exceeded:
| Surface | Budget | Why that number |
|---|---|---|
| Standing instructions | 2,400 characters | Paid in full on every turn |
| Per-turn dynamic instruction | 1,400 characters | Same, but only when it applies |
| A tool description | 400 characters | Paid every turn the tool is mounted |
| A skill description | 300 characters | A routing surface, like a tool's, but it routes to a body rather than a call |
| A skill body | 500 lines | Over that, split it |
Files that were over budget before the gate existed are listed in an acknowledgement table, each with a reason. Shrinking a file means deleting its row. A row whose file is gone is a failure. So the exceptions list can only get shorter.
Every tool has to say why it exists
Alongside the budgets there is a map that classifies every tool the agent can call into one of three kinds:
- Judgment. A model call is the capability. Take the model out and nothing is left.
- Agent-advantaged. The tool is deterministic, but a model reads the result to reason, or turns a caller's prose into its arguments.
- Pass-through. No model call, and no model gains anything. Another actor is the real caller and the model only relays arguments.
A pass-through is never shipped as a tool. It becomes a plain function, an API route, or a server action in the app that owns the screen. The test fails on a tool with no row in the map, so a new capability has to answer the question before it ships.
This is what the budgets and the map bought, measured on one caller class:
Twenty-three tools became seven for a customer, by moving groups of related actions behind one "door" tool that routes to a specialist. Every tool that stayed at the root has a written reason. The rule we used was blunt: a tool stays at the root only with a reason that cannot be refuted.
The injection that cost more than it gave
One guardrail went the other way. We keep a knowledge graph of the architecture as cross-linked markdown, and early on a hook injected semantic search results from it into every prompt, so the agent would always have the relevant design context.
It cost between 3,500 and 6,800 tokens per prompt, relevant or not. On a long session that is a meaningful share of the budget spent on context the agent did not ask for.
The hook was dropped. The instructions now say to run the search yourself before writing code, and a post-task check verifies that the graph's links still resolve. The same information, fetched on demand instead of pushed every turn. The lesson generalises: anything injected into every turn needs a per-turn cost next to it, and most things do not survive that comparison.
Decisions as records
Nineteen architecture decision records landed in the first fifteen days. That sounds like bureaucracy. In practice it is the cheapest guardrail in the repository.
An agent that starts a session does not remember last week's argument about where a schema lives. A one-page record with the decision, the alternatives and the reason means the next agent does not re-litigate it, and a reviewer can check a change against a decision rather than against taste. One of those records, "deterministic work is not agent work without an identifiable advantage", is the rule behind the tool map above.
What carried over
- Three classes, not one. Auto-fix, report, block. Most failures are not worth stopping for.
- Two incidents before a gate. And the incident is written next to the gate.
- Budgets for anything the model pays for every turn. With a test, and an exceptions list that can only shrink.
- Pull context on demand. Push nothing into every prompt that you have not priced.
- A skip needs a reason. Then it is fine.
The next part is about timing: moving verification to after the push, and what that did to how fast the work moved.