Writing

Every tool you mount is a tax on every turn

The economics of an agent's authored surface. What an instruction, a tool description and a skill each cost per turn, the budgets that enforce it, the rule that every tool needs a reason that cannot be refuted, and a manifest cut from 23 tools to 7.

An agent framework turns a directory of files into a prompt. That sentence hides a bill. Every file you add is paid for on every call the agent makes, forever, by every user. Once you see the prompt as a budget rather than a document, most decisions about where a rule or a capability belongs stop being tidiness questions and become cost questions.

23 → 7
Tools in a customer's manifest
about 3,800 to 1,900 tokens per turn
2,400 chars
Standing instruction budget
enforced by a test
400 chars
Tool description budget
500 for a door that fronts several acts
33,662
Tokens one agent sent before any conversation
before the diet; 15,800 of them tool schemas

What each surface costs

SurfaceHoldsCost
InstructionsRole, hard rules, routing logicThe whole body, every turn
A skillNew complexity, one per domainIts description every turn; its body only on the turns it loads
A tool descriptionWhen and how to call this toolThe whole text, every turn the tool is mounted
An evalWhether any of it worksNothing at runtime

Four rules follow from the table. Instructions stay thin: a rule that must always hold goes there, and anything else is paying full price every turn to be right occasionally. New complexity goes in a skill, as the default destination rather than the fallback. Repeated tool guidance belongs in the tool description, once, rather than in two skills that will drift. And depth behind a sometimes-rule goes in a skill's body, never in a new skill, because a new skill buys a shorter file and a permanently longer prompt.

One agent we measured, the one in the latency post, sent at least 33,662 input tokens on every root step before a single word of conversation: 35 tool definitions at about 15,800 tokens, most of it input schemas rather than descriptions, about 8,500 tokens of dynamic instructions, 2,800 of static ones, and 1,000 of skill descriptions. Trimming that to about 8,000 saved roughly 1.8 seconds per uncached step, and most of the step's cost. The weight was in the schemas. In most of the heaviest tools the input schema, not the description, was the bulk.

Budgets with a test

In the newest repository, the one part 2 of the guardrails series describes, the budgets are numbers in a file, and a test fails when they are exceeded.

SurfaceBudget
Standing instructions2,400 characters
Per-turn dynamic instruction1,400 characters
A tool description400 characters; 500 for a door
A skill description300 characters
A skill body500 lines, then split

A skill description is a routing surface paid every turn, like a tool's, but it routes to a body rather than to a call, so it needs far fewer words. The ceiling was set just above the longest description written before the gate existed.

Files that were over budget before the gate are listed in an acknowledgement table, each with a reason and the lane that will trim it. Shrinking a file means deleting its row. A row whose file is gone fails the test. The exceptions list can only shrink.

Every tool has to say why it exists

Alongside the budgets is a map that classifies every tool the agent can call. The question it answers is: why is a model in the loop here at all.

A pass-through is never shipped as a tool. It becomes a plain function, an API route, or a server action in the app that owns the screen. The test fails on a tool with no row in the map, so a new capability has to answer the question before it ships. The owner's version of the rule, said when there were still too many tools at the root: tools are only useful when they are called non-deterministically; we can make deterministic calls just normal functions, no?

The classification is measured, not assumed. A tool counts as making a model call only when its module graph actually reaches a delegation or an embedding call. An embedding is a lookup, and nothing reasons over the vector.

The diet, in three passes

The manifest a customer's session carried went from 23 tools and about 3,800 tokens to 7 tools and about 1,900 over a few days. The record of how is instructive because the first attempt was wrong.

Pass one: reads become dynamic. Seven read tools were mounted for everyone. Now each product's reads mount only for the classes that can use them: staff, the app's own services, and a signed-in user of that product. For everyone else they are absent, not walled off, and a one-line pointer in the instructions says where the other product's reads live.

Pass two: one door, not four. Four staff-only entry points collapsed into one tool with a field naming the desk it delegates to. A router subagent between the door and the desks was built first and rejected on measurement: two levels of delegation did not survive the framework's workflow machinery under load, the full eval pass took 131 minutes, and the credentials expired before it finished. One level of delegation was the shape proven all day, and it costs no extra model hop.

Pass three: schedule-called judgments become functions. Two moderation judgments were tools that a schedule reached by starting a session and relaying a call line through the root model. A deterministic caller with a model in the middle is the definition of a pass-through. They are now plain functions that run a prompt, a model, a typed answer and the tools it may use, with no session in front of them. The drains call them directly and assert that zero sessions start.

Then the rule that finished it: a tool stays at the root only with a reason that cannot be refuted. Nine staff-only tools became acts of their desks behind the one door. Four customer judgments became tools of one cheap-tier specialist behind a second door, one cheap model hop on a customer's turn with the per-class faces and refusals unchanged. Sixteen tools remained at the root: three doors, eight reads, four one-off actions and a search. Each of the sixteen has a reason written next to it.

Caller classDay startAfter pass twoAfter pass three
Customer3,7973,5292,662
Anonymous3,4253,1581,883
Staff4,8454,6974,082
Schema tokens2,6862,4281,274

The later passes took the customer to about 1,900 and the anonymous caller to about 1,500.

Tag every rule by how negotiable it is

A smaller surface is also a sharper one. Each rule carries its tag in its own wording: must or never means breaking it is a bug; prefer or default means deviating needs a reason; consider means take it or leave it. New rules start soft. A rule is promoted to hard only after repeated real failures, never on the strength of sounding important. If everything is hard, the model has no way to prioritise, and neither does a reader.

Layer the evals the same way

The bottom layer runs constantly: deterministic, local, blocking, on every pull request. It catches a tool that vanished, a schema that changed, a name that drifted, for free and in seconds. The surface budget test and the tool map test live there. Model-in-the-loop evals sit above, on a slower cadence, for behaviour that cannot be asserted structurally, and they should never be the first thing to go red when someone renames a file.

Restructuring an existing surface

Measure first. Description bytes per tool and per skill, and the count of references to each capability in the routing instructions. The second number decides extraction order: a capability the router mentions 21 times is not extractable until the router is rewritten, whatever its token cost. Move one domain at a time and land each with its tests. And a rule you cannot find a home for is usually two rules. Split it and route each half.

The whole discipline reduces to one habit: before adding anything to the agent's surface, say what it costs on a turn where it is not needed. Most things do not survive the question.