Green CI, six real bugs
A 328-file cleanup sweep passed every check. An extra pass asking "what could this break that no test would catch" found six real problems. The questions that found them, and why they are now a standing step.
An agent ran a sweep across a codebase: security fixes, reliability fixes, lint from 229 errors to zero with a new CI gate, framework API migrations, 24 unused packages removed, 115 components touched. Eight commits, 328 files, every check green. Types, lint, 2,474 tests, the production build.
Then the owner asked one more question: think about unintended consequences as well. It came in the same weeks we were moving verification after the push, which made the question sharper. The pass that followed found six real problems that no check had caught, and it is now a required step before any broad change is called done.
The six
Each one is a class of mistake, not a one-off. I have kept the mechanism and dropped the specifics.
An "unused" function was used by an open pull request. The sweep deleted a server action with no callers on the main branch. A pull request opened the same week, not yet merged, called it. Checking usages against main alone is checking half the repository. The pass now greps every open pull request and active branch before anything is deleted as unused.
An "unused" package was load-bearing. Two packages had no import sites anywhere. They were pinning the peer-dependency versions of a third package that was imported everywhere. Without them it resolved to a different minor version. The lockfile diff was the only place the change was visible, and nobody reads lockfile diffs. The pass now diffs resolved versions against main after any dependency removal.
New retry steps were persisting secrets. The sweep moved emails and writes inside a job runner's steps, so a retry would not repeat them. Correct. But one step returned the sign-in links it had generated, and the runner stores step return values in its run history for ninety days. A correct reliability fix created a credential store. The pass now asks, for every new step boundary: what does this return, and who keeps it.
Retries could apply stale payment state. Webhook handlers used to swallow errors and return success, so the provider never retried. The sweep made them fail properly. Now a failed event retried hours later could overwrite a newer subscription status with an older one. The fix re-reads the live state after the event and writes that instead. Every new failure path raises the same question: what retries, in what order, and what does it overwrite.
A retry storm from a shared account. The payment account is shared with another product, so the handlers see the other product's events too. Every handler was checked for whether it returned early on those or threw. One did neither on a malformed payload. A single missing guard, under the new retry behaviour, would have retried the other product's events for three days.
An availability dependency became a hard failure. Recording who owns a streaming run, so only they can re-attach to it, was written to throw on a cache error. A cache outage would have failed every chat. It now logs and refuses only the reconnect.
The seven that were fine
Half the value of the pass is the list of things that were checked and did not need to change: a narrowed query was shown equivalent to the one it replaced; a double-sanitisation was shown idempotent; a rate limit was shown to surface an error on every form that could hit it; a new browser API was confirmed unused on the client; no UI matched on database error text; the lockfile drift was all versions only a removed package had used. Each is one line in the pull request body. The reviewer reads thirteen lines and knows what was thought about.
Why green said nothing
None of the six is a bug a unit test could see, and that is the point rather than an excuse.
- Three were about other code: an open branch, a lockfile, a peer version. The test suite tests the tree it is in.
- Two were about time: a retry hours later, a stored value ninety days later. Tests run once, now.
- One was about a dependency being unavailable: a cache. Tests have the cache.
A broad change moves the ground under code that was not touched, and tests cover the code that was. It is the same gap the sweeps exist to cover from the other side. The questions that find these problems are about relationships, not units.
The pass, as a checklist
Before a broad change is called done:
- For anything deleted as unused: grep the open pull requests and active branches, not just main.
- For any dependency removed: diff the resolved versions and peer suffixes against main.
- For every new failure path, a throw where there was a swallow, a retry where there was none: what retries, in what order, what gets stored, what gets overwritten.
- For every new step or persistence boundary: what does it return, and who keeps it for how long.
- For every new external dependency on a hot path: what happens when it is down.
- For every shared resource: whose events, rows or messages can arrive here that are not ours.
- Write the list into the pull request, including what was checked and found fine.
And one more that this pull request taught: say what was not checked. This one had no browser pass, because its branch name had a slash and the preview deployment configuration excluded it. About eighty components had effects rewritten and every UI primitive changed its ref handling. Types, tests and the build passed, and none of that proves a dialog still resets or a table still paginates. The pull request said so, named the screens to click through, and left that to a human with a preview.
The rule
A green pipeline tells you the code you changed does what the tests say. It does not tell you what the code you did not change now does differently. For a change that touches hundreds of files, the second question is the one that matters, and it has to be asked on purpose, as a step, with its answers written down. We ask it now because the owner asked it once, on a day when everything was green.