Every tool you mount is a tax on every turn
The economics of an agent's authored surface. What an instruction, a tool description and a skill each cost per turn, the budgets that enforce it, the rule that every tool needs a reason that cannot be refuted, and a manifest cut from 23 tools to 7.
An agent framework turns a directory of files into a prompt. That sentence hides a bill. Every file you add is paid for on every call the agent makes, forever, by every user. Once you see the prompt as a budget rather than a document, most decisions about where a rule or a capability belongs stop being tidiness questions and become cost questions.
What each surface costs
| Surface | Holds | Cost |
|---|---|---|
| Instructions | Role, hard rules, routing logic | The whole body, every turn |
| A skill | New complexity, one per domain | Its description every turn; its body only on the turns it loads |
| A tool description | When and how to call this tool | The whole text, every turn the tool is mounted |
| An eval | Whether any of it works | Nothing at runtime |
Four rules follow from the table. Instructions stay thin: a rule that must always hold goes there, and anything else is paying full price every turn to be right occasionally. New complexity goes in a skill, as the default destination rather than the fallback. Repeated tool guidance belongs in the tool description, once, rather than in two skills that will drift. And depth behind a sometimes-rule goes in a skill's body, never in a new skill, because a new skill buys a shorter file and a permanently longer prompt.
One agent we measured, the one in the latency post, sent at least 33,662 input tokens on every root step before a single word of conversation: 35 tool definitions at about 15,800 tokens, most of it input schemas rather than descriptions, about 8,500 tokens of dynamic instructions, 2,800 of static ones, and 1,000 of skill descriptions. Trimming that to about 8,000 saved roughly 1.8 seconds per uncached step, and most of the step's cost. The weight was in the schemas. In most of the heaviest tools the input schema, not the description, was the bulk.
Budgets with a test
In the newest repository, the one part 2 of the guardrails series describes, the budgets are numbers in a file, and a test fails when they are exceeded.
| Surface | Budget |
|---|---|
| Standing instructions | 2,400 characters |
| Per-turn dynamic instruction | 1,400 characters |
| A tool description | 400 characters; 500 for a door |
| A skill description | 300 characters |
| A skill body | 500 lines, then split |
A skill description is a routing surface paid every turn, like a tool's, but it routes to a body rather than to a call, so it needs far fewer words. The ceiling was set just above the longest description written before the gate existed.
Files that were over budget before the gate are listed in an acknowledgement table, each with a reason and the lane that will trim it. Shrinking a file means deleting its row. A row whose file is gone fails the test. The exceptions list can only shrink.
Every tool has to say why it exists
Alongside the budgets is a map that classifies every tool the agent can call. The question it answers is: why is a model in the loop here at all.
- Judgment. A model call is the capability. Take the model out and nothing is left.
- Agent-advantaged. The tool is deterministic, but a model reads its result to reason, or turns a caller's prose into its arguments. Reaching it is the non-deterministic part.
- Pass-through. No model call inside, and no model gains anything. Another actor, a schedule, a route, an app, is the real caller and the model only relays arguments.
A pass-through is never shipped as a tool. It becomes a plain function, an API route, or a server action in the app that owns the screen. The test fails on a tool with no row in the map, so a new capability has to answer the question before it ships. The owner's version of the rule, said when there were still too many tools at the root: tools are only useful when they are called non-deterministically; we can make deterministic calls just normal functions, no?
The classification is measured, not assumed. A tool counts as making a model call only when its module graph actually reaches a delegation or an embedding call. An embedding is a lookup, and nothing reasons over the vector.
The diet, in three passes
The manifest a customer's session carried went from 23 tools and about 3,800 tokens to 7 tools and about 1,900 over a few days. The record of how is instructive because the first attempt was wrong.
Pass one: reads become dynamic. Seven read tools were mounted for everyone. Now each product's reads mount only for the classes that can use them: staff, the app's own services, and a signed-in user of that product. For everyone else they are absent, not walled off, and a one-line pointer in the instructions says where the other product's reads live.
Pass two: one door, not four. Four staff-only entry points collapsed into one tool with a field naming the desk it delegates to. A router subagent between the door and the desks was built first and rejected on measurement: two levels of delegation did not survive the framework's workflow machinery under load, the full eval pass took 131 minutes, and the credentials expired before it finished. One level of delegation was the shape proven all day, and it costs no extra model hop.
Pass three: schedule-called judgments become functions. Two moderation judgments were tools that a schedule reached by starting a session and relaying a call line through the root model. A deterministic caller with a model in the middle is the definition of a pass-through. They are now plain functions that run a prompt, a model, a typed answer and the tools it may use, with no session in front of them. The drains call them directly and assert that zero sessions start.
Then the rule that finished it: a tool stays at the root only with a reason that cannot be refuted. Nine staff-only tools became acts of their desks behind the one door. Four customer judgments became tools of one cheap-tier specialist behind a second door, one cheap model hop on a customer's turn with the per-class faces and refusals unchanged. Sixteen tools remained at the root: three doors, eight reads, four one-off actions and a search. Each of the sixteen has a reason written next to it.
| Caller class | Day start | After pass two | After pass three |
|---|---|---|---|
| Customer | 3,797 | 3,529 | 2,662 |
| Anonymous | 3,425 | 3,158 | 1,883 |
| Staff | 4,845 | 4,697 | 4,082 |
| Schema tokens | 2,686 | 2,428 | 1,274 |
The later passes took the customer to about 1,900 and the anonymous caller to about 1,500.
Tag every rule by how negotiable it is
A smaller surface is also a sharper one. Each rule carries its tag in its own wording: must or never means breaking it is a bug; prefer or default means deviating needs a reason; consider means take it or leave it. New rules start soft. A rule is promoted to hard only after repeated real failures, never on the strength of sounding important. If everything is hard, the model has no way to prioritise, and neither does a reader.
Layer the evals the same way
The bottom layer runs constantly: deterministic, local, blocking, on every pull request. It catches a tool that vanished, a schema that changed, a name that drifted, for free and in seconds. The surface budget test and the tool map test live there. Model-in-the-loop evals sit above, on a slower cadence, for behaviour that cannot be asserted structurally, and they should never be the first thing to go red when someone renames a file.
Restructuring an existing surface
Measure first. Description bytes per tool and per skill, and the count of references to each capability in the routing instructions. The second number decides extraction order: a capability the router mentions 21 times is not extractable until the router is rewritten, whatever its token cost. Move one domain at a time and land each with its tests. And a rule you cannot find a home for is usually two rules. Split it and route each half.
The whole discipline reduces to one habit: before adding anything to the agent's surface, say what it costs on a turn where it is not needed. Most things do not survive the question.