Read every changelog entry. Then measure it anyway.
How we upgrade a framework that releases most days. Enumerate every release between the pin and the target, classify every entry, verify the ones that matter by running them, and treat "undocumented" as a claim that needs evidence.
One of the frameworks under our agents releases most days. The other releases a security patch every few weeks. Both pin exact versions, both have to move together across repositories, and both have burned us when the bump was done by reading rather than by running. The upgrade process we settled on is a skill, a few hundred lines of instructions an agent follows, and it is the same shape for both frameworks.
Two questions, not one
An upgrade is usually framed as "does it still work". That is the smaller half. The two halves we care about are what new can we adopt, and what custom code can go. Every workaround, shim and patch we carry is a standing cost against a framework that is actively fixing things, and a bump that does not retire any of it has left value on the table.
So an upgrade is a research report first and a change second.
The steps
1. Where we are. The pin, the installed version, the newest version, and the dates of every release between them. A repository can have two gaps at once, a branch carrying a staged bump and main far behind, so the pin is checked on every branch in play and the report says which one the upgrade targets.
2. Read every release between the pin and the target. None skipped. The changelog is the primary source, but a version can appear in the changelog and not on the registry, so the list is enumerated from the changelog and cross-checked against the registry, never the reverse. The installed package's own changelog stops at the pin, so the target is downloaded into a scratch directory and read from there. Never install the target into the repository to read about it.
3. Classify every entry into one of four rows. Adopt: a new primitive a plan or a deferred item already wants, with the citation. Replaces custom code: matched against the inventory of workarounds and the patch, with the item named. Breaking: what in our code must change, found by grepping for the named API, with the files listed. Irrelevant: one clause why. Done when no entry lacks a row. Grep for the API rather than trusting the prose: a rename with one call site is a different decision from one with forty.
4. Verify by measurement, never by reading, for every adopt, replaces and breaking row. A scratch worktree, the pin bumped there, the smallest experiment that shows the behaviour: a filtered eval, a unit test, a ten-line script. When a documented claim and a measurement disagree, the measurement is what the report says, with the command.
5. Recreate the patch, if there is one. Usually the longest part. More below.
6. Write the report. What changed; the migration, measured; a re-measurement of every recorded framework fact; the workarounds this bump retires; new primitives against plans already written; the verdict. Every number cites a command.
7. Verdict. Bump now, or hold with a trigger. Hold when the newest release is under a day old and fixes nothing we have recorded, or when a breaking row costs more than the adopt rows buy. A hold becomes a deferred item with the condition that reopens it.
8. The bump lane, in its own worktree. First re-run step four on the tree it cut, because the report's file list ages against every change that landed since. Then the pin, the quarantine entry, every required peer, one commit per adoption each with its test, each retired workaround deleted together with the gate that pinned it, and the full verification at the end.
Why "measure" is the verb
The rule exists because it was broken, expensively, more than once.
The first time we filed findings upstream about the framework, two of eight entries turned out to be false; part 1 of the guardrails series has the evidence classes that came out of it. One called a behaviour undocumented when it was the documented guard, on a page in the installed package. The other said a feature had no opt-out and no docs; both existed. Entries that had been measured on a live system held up. Entries that came from reading code were where both failures happened. The rule that came out of it applies to every upgrade since: a claim that something is undocumented is itself a claim requiring evidence, and absence from a search result is far weaker evidence than absence from the pinned source.
The second kind of failure is version skew. Hosted documentation describes the latest release, and the overseer swarm learned the same lesson about a vendor's hosted search. The project pins a version. The copy of the docs shipped inside the installed package is the only one guaranteed to match what actually runs, so that is the one every load-bearing claim is checked against.
The third is that a bump had landed with no evaluation run at all. It passed types. The agent's behaviour had changed under it. Now every touched tool's evals run filtered before the bump merges, and the full pass runs at the end of the wave.
Recreating a patch
A patched dependency pins to an exact version, so every bump rebuilds the patch. Budget for it.
- Ask whether each hunk is still needed. For each one, find the behaviour upstream and check the target for it. A fix that shipped upstream is deleted, not ported. That is the "what custom code can go" half paying off.
- Reconstruct both sides. Download the old and new versions into separate scratch directories, apply the existing patch to a copy of the old one, and diff clean against patched to recover the edits.
- Port by anchoring, not by line number. Take the characters preceding each edit in the clean old file as an anchor, shrink it until it is unique in the target, insert there. Apply right to left so earlier offsets stay valid. Most hunks port mechanically.
- Re-derive when upstream restructured. An anchor that matches nowhere means the code moved. Then you read what the fix means and reimplement it at the new site. Check whether upstream absorbed part of it; it may have taken the sibling conditions and left the one that matters.
- Verify against a pristine download, never the tree you edited. One trap cost an afternoon: a patch tool run from inside any parent git repository resolves paths against that repository's root and reports "skipped" while exiting zero. Use a tool that does not, and confirm by grepping for each marker.
- Upstream it. A patch that survives a restructure is a standing cost. File the issue and record the trigger that lets the hunk be deleted.
The security case
For the web framework the shape is the same with one extra row: security. A release that names an advisory is taken now, inside the quarantine window, after opening the advisory to see whether our surface is affected. One such patch, a remote code execution in the image-generation route, was taken a day inside the window across both repositories, with the advisory id in the commit body.
The two repositories have different quarantine rules: one exempts the framework by name, the other allowlists exact versions and needs ten entries edited for a single bump. Forgetting the ten makes the install refuse the new version or silently keep the old one. The check is the installed version, not the pin.
Gotchas that cost a day each
- A build-time compile check at zero errors proves neither boot nor types. Run the type-checker beside it and boot the thing.
- A fresh worktree is not a built repository. Workspace packages that publish compiled output must be built first, or the failure looks like a framework regression.
- Peers cascade and siblings move together. Bumping a core package without its adapters leaves two copies resolved at once, and the failure surfaces as a type error far from the bump.
- A stale lockfile entry blocks the install that would replace it. Re-add the old quarantine entry, install, remove it, install again.
- A dev server started before the install still runs the old version. Restart every one.
- A configuration a framework silently ignores looks applied. Measure the behaviour, not the config.
The principle underneath
Every one of these rules is the same rule in a different coat: the framework's documentation, its changelog, its hosted search and our own notes about it are orientation, and only the pinned package running on this machine is evidence. Reading tells you where to look. Running tells you what is true. An upgrade done by reading alone is a guess with a version number.