A golden set is not an eval suite
How we keep one small labelled set per decision, score it for cents before a merge, and why the set you tuned on is not the number you report. The setup, the traps, and what the live eval suite taught us instead.
Every agent we run is full of small decisions that are really labels over text. Which specialist gets this task. Is this message a question or a request. Does this answer meet the rule. We moved most of those off free-text model calls and onto a typed judge, a small model that returns one of a fixed set of labels with a probability. That made them cheap and fast. It also made them tunable, and anything tunable needs a set to tune against.
This post is how we set those sets up, what they are for, what they are not for, and the mistakes along the way.
Two tools, two questions
A golden set is a small, labelled, version-controlled set of inputs for one decision, scored offline against the expected answer. Re-run it whenever something the decision depends on changes: a prompt, a description, a model, a tool list. It costs cents to a couple of dollars and runs in under a minute, so it runs before a merge.
An eval suite runs the whole agent end to end against a live server and asserts on what it did: which tool it called, what it said, what it spent. It answers "does the agent still behave?", costs real turns, and the owner runs it at the end of a build.
Use the set to find which decision broke. Use the suite to know the whole still works. They do not replace each other, and the most common mistake I see is using one to answer the other's question: a suite that takes four minutes and two dollars to tell you a one-line criteria change was wrong, or a set that reports 95 percent while the agent in production ignores the decision entirely.
Building one: the routing set
The worked example is the decision that gets made most often in our agents, the same kind of per-turn call the coach relies on: given a task brief, which specialist should handle it, or should the parent handle it itself. One wrong answer and the user gets the wrong tool for the job.
One JSON object per line, with a schema stated in a README beside the files:
| Field | Meaning |
|---|---|
id | Unique, prefixed by where it came from |
brief | The task as the parent would hand it over, or the user's own words for seeds |
target | The right answer: a specialist, root (the parent handles it) or clarify (ask) |
accept | Other targets that are also defensible. Strict scoring uses target only; lenient also accepts these |
kind | clear, boundary (a hard neighbour pair), vague, root, seed |
pairOf | For a vague brief, the clear brief it degrades |
source | For seeds, the test or eval case the brief came from |
The accept field is the one people skip and should not. Many routing questions have two defensible answers. A strict score that calls one of them wrong punishes the judge for a disagreement the humans also have. Reporting strict and lenient side by side tells you how much of the gap is real.
Four files, with their optimism labelled. This is the part I would copy before anything else.
| File | Cases | How it was made | How much to trust it |
|---|---|---|---|
| Development | 274 | 70 seeds mined from existing tests and eval cases, plus 220 written by hand to cover every specialist, every neighbour boundary, the themes from the agent's own failure log, root-stay cases and paired vague briefs | The criteria were tuned on this file, so its score is optimistic |
| Holdout | 74 | Written by hand after a first read of the errors | Not clean either, and the README says so |
| Holdout 2 | 124 | Generated by a model from each specialist's tool list only, never from its description | The fair number |
| Ambiguous | 20 | Context-free briefs like "Fix it." and "Send it." | Never scored as a target; only the judge's confidence is read |
The second holdout is the trick. The thing being tuned is the specialists' descriptions. A human writing holdout cases has read those descriptions and writes briefs that fit them. A model given only the tool lists has not, so its briefs test whether the descriptions describe the tools, which is the actual question.
What it reported on the development set, 218 delegated cases:
The last row is the useful one for production. When the judge was at least 90 percent confident, it was right 96.9 percent of the time, across 162 cases. That is a threshold you can route on: below it, fall back to the parent model's own judgment. On the ambiguous set, 9 of 20 briefs came back under 0.6 confidence, which is what you want from a brief that says "Fix it." The weakest slices were two neighbour boundaries, where one kind of request kept landing on the specialist next door. The set names them; the confusion matrix is in the output. That is what "which decision broke" looks like.
And the baseline it was measured against. The same briefs were given to the parent model with its real tool list, recording which tool it would call without running it. It delegated when it should 81.7 percent of the time, picked the right specialist 97.8 percent of the time once it delegated, and landed at 82.6 percent end to end. So the judge was better at deciding whether to delegate, and the model was better at picking whom. The study that produced this set concluded the judge-as-router pattern was not worth adopting. The set is what was left worth keeping.
The rules the set lives under
- Beside the code, with a README. Schema, how each file was made, how to re-run, what it costs, when to re-run. The README is where the optimism labels live.
- A tuning set and a holdout you never tune on. When the holdout gets read for errors, it stops being a holdout. Generate a new one.
- No real user data. Briefs use invented names and reserved example hosts. A unit test reads every set file and fails on an email address, a handle, a non-example URL or a phone number, using the same patterns the product uses to detect user details. It runs on every commit.
- The set stays in sync with the agent. The same test checks that every target names a specialist directory that exists and every root tool names a file that exists. When a specialist was folded into the parent, its cases kept their ids and briefs and were relabelled, and the README records the mapping.
- On demand, never in CI or a hook. It spends money. A package script runs it; the latest baseline and its date go in a central table of sets.
- Compare against something. Two cents of judge versus two dollars of model baseline. Without the baseline, 90 percent is just a number.
Golden sets from history
The routing set was written. The other kind is exported.
When a judgment has a human-decided history, the golden set comes from the outcomes, not from the old prompt text. For an answer-selection capability, the set is 200 closed questions with the selections humans finally made, read from the production database through a role that can only SELECT, by a human running one script, because no agent has production access. The exit bar for the new judge was agreement of 90 percent or better with the human-final selections, and the plan noted the risk plainly: the judge is only as good as the golden set, so sample across sources and across answer counts.
The export redacts as the last step before the write, so nothing can be added after it. Email addresses and links become tokens. Author identifiers are hashed with a salt per export, so two exported files cannot be joined back into one identity. Job titles are reduced to a role token. The answer text itself stays whole, because it is published content and the judge has to read it to be worth judging. Each of those rules has a unit test.
A third flavour is the byte-for-byte golden: fifteen email templates, each with its rendered HTML and text saved under two time zones. A 250-line refactor across all fifteen landed as a diff of zero against the goldens. Without them it would have been out of scope for the lane that did it.
Auditing a live judgment
Once a judgment is in production, the golden set stops being the only evidence. We audit against real rows, with a recipe that has held up across 34 live judgments:
- Pull up to 50 real rows the judgment actually scored.
- Rescore them with the current criteria wording, pulled from source, not from memory.
- Hand-review every disagreement.
- Try wording variants on the same rows. Zero regressions allowed. Leave the thresholds alone.
That last instruction came from the finding that surprised me most. In every underperforming judgment, the errors sat in a flat band between 0.35 and 0.65 confidence. Naming the observed failure shape in the criteria fixed them. Moving the threshold never did. A threshold tunes how often you are wrong at the margin; the criteria decide where the margin is.
One trap for anyone batching a judge: a batch call that hits a rate limit can return short and silently misalign every later row with the wrong input. Chunk small and assert that the output count matches, or score one row per call.
What the eval suite taught us that sets cannot
The live suite is where the expensive lessons were, and they are worth the two dollars.
Provenance classes. Every case in our suite is labelled either incident-derived, meaning we once got this wrong and no longer do, or rule-derived, meaning we said we would not and still do not. A green incident-derived case is regression evidence. A green rule-derived case proves only our stated intent; it cannot tell you the failure was ever possible. Cases invented from what the rules say are worth much less than cases captured from what actually went wrong, and the label sits at the definition site so nobody counts them the same.
Contamination. The harness ran every case in one invocation under one conversation id, and the agent bootstraps a conversation's history into each new session. So case N saw cases 1 to N−1. One case failed in the full run and passed alone, a false red: the agent was obeying a rule not to re-run a search whose results were already in the thread. Another passed in the full run and failed alone, a false green over the suite's one known live defect. Every full-suite number before that day had been measured under contamination. One case per invocation, a fresh conversation per run, and the spread across three identical runs went from about six points to exactly zero on every case.
Mutation testing against the rule. For two rule-derived cases we deleted the rule they quoted, then inverted it to command the opposite, and ran the case after each mutation. Both stayed green through all four. The behaviour came from the base model, and our instruction layer could not move it in either direction. The cases were kept and relabelled as model-swap guards. The doctrine generalises: a rule-derived case has to be mutation-tested against the rule it quotes, or you cannot tell whether you are measuring your instructions or your model vendor.
The instrument before the measurement. Three independent defects, found within a day of anyone first running the suite, each made a gate report coverage it did not have. The question for every new case became: has this assertion ever been observed to fail for the right reason? And for every new helper: can this matcher fail at all, on an agent that produced no output?
When not to
For a product whose shape was changing weekly, we stopped writing and extending eval suites for a while. The owner's words: we are too early for evals, we do not even know what the product is. What survived that pause was the cheap layer: golden sets of a few dozen cases, and throwaway judge batches of eight or ten examples to pick a threshold. Suites come back when the product settles. Sets never left, because they cost nothing to keep and cents to run.
How to set one up
- Pick one decision. A label, a score or a yes/no over text. If it needs the whole agent to evaluate, it is a suite case, not a set case.
- Write the schema down, including an
acceptfield for defensible alternatives, and score strict and lenient. - Mine seeds from what already exists, then write cases for every class, every neighbour boundary and a vague twin of each clear brief.
- Make a holdout you cannot have been influenced by. Generate it from the inputs the criteria are supposed to describe, not from the criteria.
- Label the optimism of every file in the README.
- Enforce no real data with a test, and keep the set in sync with the code with the same test.
- Run on demand, record the baseline and its date, and compare against the model you are replacing.
- Once live, audit on real rows. Fix the criteria, not the threshold.
If you keep golden sets a different way, or have a better answer to the holdout problem than "have a model write it from the tool list", I want to hear it.