Writing

Nobody works past their context limit

How a long migration was run by a chain of fresh-context worker agents, supervised by a chain of fresh-context overseers, with a cold review before anything touching money shipped. The design, the gates that caught real defects, and the honest costs.

The default way to run a long agent task is one long session. It works until it does not: the context fills, the quality degrades without announcing itself, and at the end the same agent that wrote the code reviews it and ships. We ran a multi-week migration a different way, and the document that came out of it is the most useful thing I own about running agents at scale. This is the short version, with the numbers.

~200K tokens
Worker cycle
about 23 minutes, then a hard stop
9 of 12
Cold reviews that returned HOLD
right every time; four found a billing defect
~180K tokens
Cost of one cold review
about 12 minutes
677 KB
Tracker every cycle read first
46% of it was closed items

In three sentences

A chain of worker cycles each start with fresh context, do one wave of work, and hand off through a file before they run out. A chain of overseer generations does the same thing one level up, holding continuity in a file rather than a context. Nobody works past their context limit, and no code on a money path ships without a cold review by an agent that did not write it.

The overseer never writes production code. Its job is memory, arbitration, and refusing to let a tired agent push.

Set up before the first worker starts

Skipping any of these produces a swarm that argues with itself.

  1. Three law documents in the repo, referenced by every brief: the playbooks with numbered gates; the scope map, where every unit of work is a numbered item with a verdict; and a tracker that holds current state only, updated in the same commit as the change it describes, paired from day one with an archive for everything closed.
  2. A fixed handoff path, one per lane, overwritten each cycle. The only thing a new cycle reads first.
  3. An overseer handoff in the repo, not in a temp directory, because it outlives more than one cycle and carries the standing constraints.
  4. A standing-law block in the agent instructions, so every agent loads the non-negotiables without being told.
  5. A verification gate that is cheap to state and expensive to fake. Ours was the full uncached test and type-check run, with the task count recorded as the coverage signal so a silently dropped suite is visible. The tally line is pasted from the run, never retyped, the rule part 1 of the guardrails series tells the longer story of.
  6. A branch policy. One long-lived feature branch, never commit into an open pull request's branch, nothing red ever lands on it.
  7. A parity checklist if you are replacing something that works. Progress bars measure construction. Only a parity list measures whether shipping would be a regression.

The overseer was the single point of failure

Every worker could die safely; that was the design. The session supervising them held the standing constraints, the decision rules and the reasons behind both in nothing but its own context. Three fixes were tried.

The first was a file plus a human typing a reset at the right moment. Correct about the file, wrong about the trigger: a mechanism that needs a human to act on time is a chore with a deadline. The second split the overseer into a thin persistent dispatcher and disposable judgement cycles, a real improvement that postpones the problem rather than removing it, because the dispatcher fills eventually. The third made the overseer a chain, exactly like the workers.

That rested on a distinction that is easy to miss. A subagent is a child: it dies when its spawner ends. A background session is a peer: it has its own process and outlives whoever started it. The rule "anything spawned by a bounded agent is bounded by it" is true of children and false of peers, and the whole handover depends on knowing which you have. The handover is: the outgoing overseer rewrites the handoff file, spawns its successor as a peer, sends it the brief, and tells every worker who their new boss is. No human trigger.

One landmine, recorded because it is the system's own rule turned on its author: spawning a background session with the prompt on the command line silently dropped the prompt and returned success, leaving a child idle forever. A failure must be distinguishable from success, not merely closed.

The gates that stopped real defects

A verdict that has not returned is treated exactly like a HOLD. A verdict taken against an older tree is stale and does not count.

The review nobody runs

Every piece of code got a cold review by an agent that did not write it. The overseer got none. It wrote the rules, judged compliance with the rules and decided what the run did next, the one role exempt from the system's central discipline.

Over one day, almost every structural improvement to the run came from the human, not the overseer: that a local database made per-lane databases free, that agents had grep the whole time and never needed to load an 80K-token file, that documents could be arranged by load decision, that two user-facing surfaces were regressions rather than unbuilt features. The overseer produced good analysis after each prompt and initiated none of them. That is what an unreviewed role looks like from the inside: competent execution, invisible blind spots, and a human quietly doing the review job by hand.

The fix is a standing process critic with fresh context and no stake in prior decisions, briefed against the overseer's own interest: what is measured that does not matter, which stated rules are not actually followed, where is risk accumulating unnamed, what would an experienced engineer call obviously wrong that everyone stopped seeing thirty cycles ago. Cap it at five findings and act on all five. And a human should read its output rather than the overseer's summary of it.

Honest costs

UnitTokensWall-clock
Worker cycle200–210K~23 min
Cold review~180K~12 min
Landing one lane~127K~11 min
Search fan-out on a cheap model~74K~4 min
Full uncached gatenot measured in tokens2m09s to 4m46s

The gate row is a range on purpose. Two uncached runs an hour apart, no code change, differed by 1.8 times. The task count was stable; the wall-clock was dominated by how warm the machine was. Quote a single figure and you imply it is repeatable.

Rework dominated. Reviews plus the fix waves they triggered came to about half the total spend. The reviews were right every time, so the lesson is not to review less. It is that anything a compiler could have asked is enormously cheaper than anything a reviewer has to find, and cheaper still than what a reviewer cannot find at all. One defect survived two reviews as an if-ladder and died instantly when the same logic became a satisfies-bound table with a total classifier.

The orientation tax was the largest recoverable cost, and it was invisible. The tracker reached 677 KB that every cycle read before doing any work. Measured in lines, the closed-items section was under one percent of the file. Measured in bytes it was 46 percent: 44 rows whose longest were 7.6 KB each. The archival pass that followed a line-based table moved what the table pointed at and left the single largest item untouched. Count in the unit that carries the cost.

The fix generalises: split by kind, not by age. State should be generated, never hand-written. Evidence belongs in an archive with a pointer inline. Decisions, the rule and its why, stay in the briefing document, uncompressed, because the why is exactly what a successor cannot re-derive. Prose has no definition site, which is why we stopped writing comments. The tempting move is to compress the rules for agents. The evidence pointed the other way: a terse one-line rule was violated three times and only stuck once it carried its reasoning.

Depth is the failure mode you will not notice. Six consecutive cycles and about 1.8 million tokens went into one slice while the main branch did not move, every cycle finding a real defect, each narrower than the last, half of them in code the same run had just written. Reviews finding real problems is not by itself evidence that reviewing again is the best next move. Watch whether the branch moves, not whether the findings are valid.

The shortest version

Write the law down first. Give every agent a fresh context and a hard stop, including the overseer, whose successor is spawned as a peer, told what it owns and handed a file. Never let the author of a change be its reviewer, and that includes the overseer: spawn a standing critic of how the run is managed, or a human ends up doing that job by hand and you will not notice. Measure parity, not just construction. Verify state yourself; a command that exits zero has not told you it did anything. Park early, but watch that parking does not become the work. Record your own errors as yours, because the next generation reads that record and believes it.