Delete the comments. The agent believes them.
Code comments were a courtesy to the next human. Coding agents read them as fact, write them as narration, and both go wrong. How we got to a zero-comment budget on new code, and what we do with the context instead.
I used to think comments were free. They cost nothing to run, they help the next reader, and the worst case is that one goes a little stale. Then most of the readers of our code became coding agents, and the worst case changed.
Comments are read as code
A human reader treats a comment as a hint, weighted by how old the file looks. An agent treats it as an instruction from the author, with the same authority as the code beside it, and it has no way to know the file is old. Three cases from our repositories, each of which cost real time.
The comment that recorded a rejected decision. A component file carried a header comment explaining that an always-visible row of filter chips had once been rejected because it showed internals to the end user. The comment was accurate when written. Weeks later the objection no longer held, the owner said so, and the agent kept citing the comment as a reason not to build visible filters. The code had moved on; the comment was still issuing a verdict. We had to write a separate note telling the agent the comment was stale, which is a strange thing to need.
The docblock that lied about behaviour. A hook on the agent's shell tool said, in its own docblock, that it fired on commit and push. What it actually gated on was whether anything was staged, which with a dirty index is always. So it ran before every shell command, including listing a directory. Measured on a large merge: 100 seconds per command, about 25 commands in a session. Nobody caught it for a while, partly because the comment said it was fine. I wrote the longer story in part 1 of the guardrails series.
The docblock that described someone else's code. A test file carried a 55-line comment describing the internals of a framework dependency at one version: three call sites, all unguarded. By a later version the framework had guarded one of them. The comment still asserted all three. It was read and repeated as fact by an agent before anyone checked the shipped code. The conclusion happened to survive; the reasoning did not.
The pattern across the three: a comment is prose with no definition site. Nothing keeps it in sync with the thing it describes, nothing fails when it drifts, and the reader that now dominates has no instinct for staleness.
Comments are written as narration
The other half is what agents write. Left to themselves they comment generously: restating the signature, narrating the diff ("added a null check here"), retelling how they debugged something, dating the change. That text is useful exactly once, to the person reviewing that commit, and it belongs in the commit message, where it is attached to the change and leaves when the change is superseded.
We tried a budget first. The instruction said a comment should be a why, a non-obvious constraint, or a trap someone would reintroduce; one line, rarely three. The repository's checker enforced three lines as a ceiling. What happened is worth recording: a ceiling reads as a target. New code arrived with three-line blocks on nearly every declaration and passed the gate. The rule was satisfied and the intent, default to none, was not.
The deeper cause is style contagion. Anyone editing a file with twenty-line docblocks, human or model, matches the density in front of them, because matching the surrounding style is otherwise correct behaviour. A written instruction loses that argument every time, because the file is the more immediate signal. A gate does not lose it.
The rule we landed on
In the main repository, the standards checker now enforces this at commit time, at error severity:
- No comments on lines you added. None. If the code needs explaining, rename or extract until it does not. Context worth keeping goes in a markdown document or the commit message, where it can be read, updated and reviewed as prose instead of drifting inside a source file.
- Directives are exempt, because a tool reads them and behaviour changes if they go: lint disables, type-checker pragmas, coverage ignores, reference directives, licence headers. These are code that happens to look like a comment.
- Never a comment about code you do not own. A dependency path, a pinned version, "the behaviour changed in 0.36". It goes stale on the next bump and then misleads with authority. Assert the behaviour in a test instead. A test that executes cannot go stale silently; it fails.
- Judged on added lines only. The repository carries large legacy docblocks. A whole-file rule would bury every commit in findings about code nobody touched, so the checker reads the diff. Legacy blocks are invisible until you edit them, and then the house rule is: trim what you touch.
In the other repositories, the instruction file carries the softer version, default to none with a why allowed, plus "prefer a rename or an extraction over a comment". The difference is honest. The strict rule exists where the agents could not reliably tell a why from a narration, and the checker exists because the instruction alone did not hold.
Where the context goes
The ban is on source files, not on documentation. The objection to zero comments is usually "but some things need explaining", and they do. The question is where the explanation can be kept correct.
| What you wanted to write | Where it goes now | Why there |
|---|---|---|
| Why this is done this way | The commit message, or a decision record | Attached to the change; reviewed as prose; does not drift inside the file |
| A constraint from a dependency | A test that exercises it | Fails when the dependency changes instead of lying |
| A trap someone would reintroduce | A checker rule | Blocks the reintroduction instead of hoping it is read |
| How I debugged this | The pull request | Useful once, to the reviewer |
| A rejected design decision | A memory or a doc with a date | Can be marked stale when the decision changes |
One case made the last row concrete. A test failure that only appeared in the full suite and passed in isolation was written up, twice, as a flaky gotcha, and "fixed" twice against the wrong cause before anyone looked at the fixtures. The fix was not a comment at the site explaining the trap. It was a checker rule that refuses the pattern, with the explanation of why living in the checker, which is where someone hitting the rule will read it.
The leftovers
New comments are easy to stop at the gate. Old wrong comments are the harder problem, because they sit in files nobody is touching, and each one is a small lie waiting for an agent to read it.
So the weekly scheduled run sweeps them. Its scope is deliberate: wrong comments are leftovers from previous releases, not in new pull requests. For each change merged that week, it lists what changed in behaviour, renamed symbols, changed defaults and limits, and then searches the whole repository for comments and docs that still describe the old behaviour, outside the change's own diff. It adds one rotating directory a week for comments that contradict the code beside them. Each candidate gets a second opinion from a cheap classifier, "does this comment contradict or merely restate the code shown", and a split verdict means leave it and list it. About twenty files per pull request, no behaviour changes mixed in, and when the comment and the code disagree it checks the history for which one is the bug before deleting anything.
The same sweep removes tests that only restate the code they check, for the same reason: both are prose about the code pretending to be a guarantee.
What I would tell someone starting
- Assume the agent believes every comment. Then read your codebase's comments as instructions and ask how many are still true.
- Set the budget for new comments to zero, exempt directives, and enforce it at commit on added lines. A ceiling becomes a target; zero has no target.
- Move the context. Commit messages, decision records, tests that execute, checker rules. Prose about code goes where prose is reviewed.
- Sweep the leftovers on a schedule, scoped to the ripple of each week's changes, with a second opinion before any deletion.
- Keep the rule for humans too. Comment density in our oldest repository is 4.2 percent of lines; in the newest it is under 1 percent. The newest is the one agents navigate best.