Stop growing the system prompt. Coach the agent instead.
We moved behaviour rules out of the system prompt and into one-line, per-turn nudges picked by a cheap judge. Here is what we measured.
We run an eve agent in production at Featured. For a while, every time it misbehaved in some specific situation, the fix was another rule in the system prompt. Each rule fixed the one thing and quietly regressed something else, and the prompt kept growing.
Most of those rules only matter on some turns. Paying for them on every turn, and risking collisions between them on every turn, felt wrong. So yesterday I wrote up what we did instead and posted it to the eve discussions. This is the longer version.
The idea
A cheap judge model reads the incoming message, or the turn that just finished. When a rule applies, the agent gets one short line of guidance for that turn only. When no rule applies, the agent pays nothing.
We started calling it a coach. It does not rewrite answers and it does not block them. It says "do X this time" and gets out of the way.
What we measured
1. Does a one-line nudge even reach the model?
Before measuring anything real, we checked that a single line appended to a turn is actually read, across multi-turn conversations. We used a nonsense rule, "always include the word banana", so there was no baseline behaviour to confound it.
| Placement | Turns that complied |
|---|---|
| No nudge (control) | 0/8 |
| Appended as user-role context for the turn | 8/8 |
| System-scope instruction for the turn | 8/8 |
| Client-supplied ephemeral context on send | 8/8 |
Every placement works. That left us free to choose placement on cost, which turned out to matter a lot.
2. Does it change a real behaviour?
On the first turn of a new conversation, our agent often reached for generic web search when the person wanted our own domain search. We sent a one-line nudge, only when a judge said the message was that kind of ask.
Run 2 used a fuller search index, which is why the control improved slightly. The nudged runs were perfect both times.
3. What does it cost in prompt caching?
This decided the placement for us. Prompt caching is where most of an agent turn's latency lives. Putting per-turn content into system scope dropped prefix reuse on the first model call of that turn from roughly 94 to 97 percent down to 0 percent. It recovered on later calls within the turn, but that first call is the one the person is waiting on.
Appending the nudge as ephemeral context at the end of the conversation kept reuse intact.
| Placement of the nudge | Prefix reuse on the turn's first call |
|---|---|
| None | 94–97% |
| System scope | 0% |
| Ephemeral context, end of transcript | 94–97% |
4. Can a judge decide when to coach, cheaply and safely?
One batched classification call per message, made before the turn starts, on a typed judge. We set a high threshold because a wrong nudge costs more than a missed one. A miss just falls back to default behaviour.
| Set | Labels correct | False nudges | Misses | Latency p50 / p90 |
|---|---|---|---|---|
| Tuning (60 cases) | 57–58/60 | 0 | 4–5 | 270 ms / 470 ms |
| Fresh holdout, run once (22 cases) | 18/22 | 1 | 4 | 277 ms / 455 ms |
Under 300 milliseconds at the median, and the errors lean the safe way.
5. Instructions alone don't always hold
One prompt rule, "offer choices rather than ending on a question in prose", was broken in 10 of 10 sampled answers. It was in the system prompt the whole time.
That is the kind of rule we are now moving to a deterministic check after the turn, with a nudge on the next turn only when it was broken.
What didn't work
Not everything works. One nudge we tested was clearly received and then ignored. I don't have a tidy explanation for it yet.
The lesson I took: each nudge has to earn its place with its own measurement. "We added a rule" means nothing until you have counted how often it fired and whether the next answer complied.
What first-class support could look like
If this pattern belongs in the framework rather than bolted on beside it, here is what I would want.
- Pre-turn coaching. A hook that sees the incoming message, and optionally the recent transcript, and returns guidance that applies to this turn only. Ephemeral, never written into history, placed where it cannot break the prompt cache. Today instruction resolvers at turn start can't see the incoming message, so pre-turn coaching has to happen outside the agent or in the client.
- Post-turn coaching. A check after the turn that can leave a pending note, delivered once on the next turn and then cleared.
- Measurement built in. For every coaching rule: how often it fired, and whether the next answer complied. Without that, nudges turn into hope.
- Budgets and sampling. Cheap deterministic checks every turn, judge calls sampled for low-risk rules, and per-rule thresholds that can be tuned without a deploy.
Open questions
I asked the eve team four things, and I would ask anyone running agents the same:
- Is this a framework primitive, or are instructions plus client context the intended way?
- Should pre-turn hooks be able to see the incoming message?
- Would you guarantee a cache-safe placement for per-turn context?
- Is anyone else running judge-in-the-loop guidance like this, and what did you measure?
If you have numbers, I want to see them. The discussion is the best place to reply.