Guardrails for coding agents · Part 4
The agent made seven decisions without me. I reverted all of them.
A week of regressions from choices the agent made on its own, and the rules on the human side of the loop that came out of it.
Everything in part 3 made the loop faster. This part is about the week when fast turned into damage, and about the guardrails that matter most: not the ones on the code, but the ones on the conversation.
What happened
We had a working, fast app. Over a few days of parallel agent work it got slower, then started breaking in small, compounding ways. Tracing it back, every cause was a decision the agent had made without asking: a new routing structure, a layout that always split the screen, a redesigned URL scheme, a change to how the agent handed off between views, and, on my own account, archiving my data and switching me to a test workspace.
Each was defensible on its own. None had been put to me. My sentence at the end of that day: the problem is you made decisions without sign-off.
Four things in how the agent handled the aftermath each became a rule:
- It blamed infrastructure latency before measuring. The cause was in code it had written.
- It pushed a build that passed type-checking and failed on the hosting platform. A type check is not a build.
- It kept patching forward on top of the redesign instead of returning to the last commit known to be good.
- It described its own regression as the established pattern. A week earlier it had moved first-paint data reads from the server to the client, 140 reads across 53 files; the latency post has the numbers. Asked why the app was slow, it defended that as how the codebase worked.
Measure, never assert
The first rule: state a cause only with a measurement attached. We wrote a small script that logs first byte, time to the composer, time to content and every slow request against any origin. Every performance claim since has arrived with its output pasted in.
The second: when the owner says it was working before, find the last good commit first. Older deployments can be re-aliased to the test environment in seconds, so "restore the last good build, then measure" costs almost nothing. Debugging forward from a broken state is the expensive path, and it is the agent's default.
The third is the push gate, which replaced "type-check passed" as the definition of done:
- type checks across the touched packages,
- a local production build, because the failed deploy had passed type checks,
- a browser pass over seven named flows, from the home page through the main views to sign-out,
- then confirm the deployment succeeded.
Any failure blocks the push. It is slower than before, and much faster than a day of finding out.
A ledger, and no delegation
The owner's instruction, close to verbatim: make sure you don't miss any of my instructions or thoughts, work off a ledger, no delegation.
So the session kept a ledger. Every instruction or stray thought became a numbered row the moment it arrived, with a status: open, doing, done with a commit hash as proof, or waiting on a named person or process. Work proceeded top-down from it. By the end of the week it had over eighty rows, and "what is still pending from our original list" was answered by reading, not remembering.
The ledger also carries standing rules at the top, the ones that must survive context compaction: the last good commit and why, the push gate, measure never assert, sign-off first, and "communicate before acting on anything surprising, never go quiet".
No delegation was a step back from part 3's parallelism, and deliberately so. For a week, the cost of agents making calls I could not see outweighed the throughput. Delegation came back later, for background proposal work, one item at a time.
Which decisions are mine
The rule that came out of the week is a classification, not a blanket "ask first". A blanket rule would have put us straight back to the five-minute-pause problem.
| The agent decides | I decide |
|---|---|
| How to implement an approved change | What the product looks like or does |
| A bug fix that restores already-approved behaviour, and says so | Routing and layout structure |
| UI decisions and new functionality once the gate is green | What data is written, archived or moved |
| Which option to take when the owner has said "don't block on me" | What the agent says to users |
| Every recommendation in a sweep it agrees with, not a subset | Merging to main, and anything that touches production |
Two refinements landed the same week.
Options as mocks, never as questions. Asked to choose a layout in words, I could not. My line: show me options to pick from, I cannot take a call without seeing it first. So a design decision now arrives as mocked screens, screenshotted, with a recommendation. Do not ask whether to mock. Always mock. And the mock harness has to cover every state of a surface, not just the happy path, so the empty tab and the overflowing list are caught on the mock rather than one at a time on the live environment.
Self-critique before showing anything. After I found a set of mocks broken in ways the agent would have seen had it looked, the rule became: it should be mistake-free when you show me. Every screen, not a sample, at three widths, both themes, driven like a first-time user. Data has to agree with itself between the list and the card. No developer notes in the copy.
The pendulum
The clean table above hides the shape of the week, so here it is.
On Tuesday the rule was sign-off on everything, with work held until I answered a keep-or-revert list of seven items. By Friday, deep in a build, the instruction was the opposite: don't block on any more inputs, give me a report at the end, prioritise the better product and the better UX.
Both were right for their day. What reconciles them is that Friday's rule came with hard limits that were not questions: no production changes, the shared test environment hands-off, a spend cap that stops the work rather than asking. Inside those, decide and report. Outside them, there is nothing to decide, so there is nothing to ask.
One line from mid-week is the rule I keep: no existing practice is sacred, only our goals are, good UX and fast UX. Followed minutes later by: but things that have proven to work need a really good reason to be cast aside. The 140 client reads failed that test. The new routing failed it. Neither had a stated reason, let alone a measured one.
What carried over
- A measurement with every claim about cause.
- Last good commit first. Restore, then debug.
- The push gate. Types, a production build, a browser pass, deploy confirmed.
- A ledger with standing rules at the top. Every ask is a row.
- A decision table, not a blanket rule. Product, structure, data, the agent's voice and production are the human's.
- Mocks instead of questions. Self-critique before showing.
- Hard limits are not questions. Inside them, decide and report.
The last part is about the one rule in this series that was written down, repeated, and still did not hold, and what finally did.