Writing

Delete the comments. The agent believes them.

Code comments were a courtesy to the next human. Coding agents read them as fact, write them as narration, and both go wrong. How we got to a zero-comment budget on new code, and what we do with the context instead.

I used to think comments were free. They cost nothing to run, they help the next reader, and the worst case is that one goes a little stale. Then most of the readers of our code became coding agents, and the worst case changed.

0
Comment budget on new lines
enforced at commit, directives exempt
55
Lines of one docblock about a dependency's internals
wrong by the next version
100 s
Cost of one docblock that lied
per shell command, around 25 commands a session
4.2%, 3.0%, 0.9%
Comment lines in three repos
of TypeScript lines; the oldest repo is the highest

Comments are read as code

A human reader treats a comment as a hint, weighted by how old the file looks. An agent treats it as an instruction from the author, with the same authority as the code beside it, and it has no way to know the file is old. Three cases from our repositories, each of which cost real time.

The comment that recorded a rejected decision. A component file carried a header comment explaining that an always-visible row of filter chips had once been rejected because it showed internals to the end user. The comment was accurate when written. Weeks later the objection no longer held, the owner said so, and the agent kept citing the comment as a reason not to build visible filters. The code had moved on; the comment was still issuing a verdict. We had to write a separate note telling the agent the comment was stale, which is a strange thing to need.

The docblock that lied about behaviour. A hook on the agent's shell tool said, in its own docblock, that it fired on commit and push. What it actually gated on was whether anything was staged, which with a dirty index is always. So it ran before every shell command, including listing a directory. Measured on a large merge: 100 seconds per command, about 25 commands in a session. Nobody caught it for a while, partly because the comment said it was fine. I wrote the longer story in part 1 of the guardrails series.

The docblock that described someone else's code. A test file carried a 55-line comment describing the internals of a framework dependency at one version: three call sites, all unguarded. By a later version the framework had guarded one of them. The comment still asserted all three. It was read and repeated as fact by an agent before anyone checked the shipped code. The conclusion happened to survive; the reasoning did not.

The pattern across the three: a comment is prose with no definition site. Nothing keeps it in sync with the thing it describes, nothing fails when it drifts, and the reader that now dominates has no instinct for staleness.

Comments are written as narration

The other half is what agents write. Left to themselves they comment generously: restating the signature, narrating the diff ("added a null check here"), retelling how they debugged something, dating the change. That text is useful exactly once, to the person reviewing that commit, and it belongs in the commit message, where it is attached to the change and leaves when the change is superseded.

We tried a budget first. The instruction said a comment should be a why, a non-obvious constraint, or a trap someone would reintroduce; one line, rarely three. The repository's checker enforced three lines as a ceiling. What happened is worth recording: a ceiling reads as a target. New code arrived with three-line blocks on nearly every declaration and passed the gate. The rule was satisfied and the intent, default to none, was not.

The deeper cause is style contagion. Anyone editing a file with twenty-line docblocks, human or model, matches the density in front of them, because matching the surrounding style is otherwise correct behaviour. A written instruction loses that argument every time, because the file is the more immediate signal. A gate does not lose it.

The rule we landed on

In the main repository, the standards checker now enforces this at commit time, at error severity:

In the other repositories, the instruction file carries the softer version, default to none with a why allowed, plus "prefer a rename or an extraction over a comment". The difference is honest. The strict rule exists where the agents could not reliably tell a why from a narration, and the checker exists because the instruction alone did not hold.

Where the context goes

The ban is on source files, not on documentation. The objection to zero comments is usually "but some things need explaining", and they do. The question is where the explanation can be kept correct.

What you wanted to writeWhere it goes nowWhy there
Why this is done this wayThe commit message, or a decision recordAttached to the change; reviewed as prose; does not drift inside the file
A constraint from a dependencyA test that exercises itFails when the dependency changes instead of lying
A trap someone would reintroduceA checker ruleBlocks the reintroduction instead of hoping it is read
How I debugged thisThe pull requestUseful once, to the reviewer
A rejected design decisionA memory or a doc with a dateCan be marked stale when the decision changes

One case made the last row concrete. A test failure that only appeared in the full suite and passed in isolation was written up, twice, as a flaky gotcha, and "fixed" twice against the wrong cause before anyone looked at the fixtures. The fix was not a comment at the site explaining the trap. It was a checker rule that refuses the pattern, with the explanation of why living in the checker, which is where someone hitting the rule will read it.

The leftovers

New comments are easy to stop at the gate. Old wrong comments are the harder problem, because they sit in files nobody is touching, and each one is a small lie waiting for an agent to read it.

So the weekly scheduled run sweeps them. Its scope is deliberate: wrong comments are leftovers from previous releases, not in new pull requests. For each change merged that week, it lists what changed in behaviour, renamed symbols, changed defaults and limits, and then searches the whole repository for comments and docs that still describe the old behaviour, outside the change's own diff. It adds one rotating directory a week for comments that contradict the code beside them. Each candidate gets a second opinion from a cheap classifier, "does this comment contradict or merely restate the code shown", and a split verdict means leave it and list it. About twenty files per pull request, no behaviour changes mixed in, and when the comment and the code disagree it checks the history for which one is the bug before deleting anything.

The same sweep removes tests that only restate the code they check, for the same reason: both are prose about the code pretending to be a guarantee.

What I would tell someone starting

  1. Assume the agent believes every comment. Then read your codebase's comments as instructions and ask how many are still true.
  2. Set the budget for new comments to zero, exempt directives, and enforce it at commit on added lines. A ceiling becomes a target; zero has no target.
  3. Move the context. Commit messages, decision records, tests that execute, checker rules. Prose about code goes where prose is reviewed.
  4. Sweep the leftovers on a schedule, scoped to the ripple of each week's changes, with a second opinion before any deletion.
  5. Keep the rule for humans too. Comment density in our oldest repository is 4.2 percent of lines; in the newest it is under 1 percent. The newest is the one agents navigate best.