Writing

A golden set is not an eval suite

How we keep one small labelled set per decision, score it for cents before a merge, and why the set you tuned on is not the number you report. The setup, the traps, and what the live eval suite taught us instead.

Every agent we run is full of small decisions that are really labels over text. Which specialist gets this task. Is this message a question or a request. Does this answer meet the rule. We moved most of those off free-text model calls and onto a typed judge, a small model that returns one of a fixed set of labels with a probability. That made them cheap and fast. It also made them tunable, and anything tunable needs a set to tune against.

This post is how we set those sets up, what they are for, what they are not for, and the mistakes along the way.

2 ¢
Cost of one run
about 30 seconds, for one decision
90.4%
Accuracy on the set we tuned on
not the number we report
94.0%
Accuracy on the set we never saw
the number we report
0
Real user messages in any set
a test fails the commit otherwise

Two tools, two questions

A golden set is a small, labelled, version-controlled set of inputs for one decision, scored offline against the expected answer. Re-run it whenever something the decision depends on changes: a prompt, a description, a model, a tool list. It costs cents to a couple of dollars and runs in under a minute, so it runs before a merge.

An eval suite runs the whole agent end to end against a live server and asserts on what it did: which tool it called, what it said, what it spent. It answers "does the agent still behave?", costs real turns, and the owner runs it at the end of a build.

Use the set to find which decision broke. Use the suite to know the whole still works. They do not replace each other, and the most common mistake I see is using one to answer the other's question: a suite that takes four minutes and two dollars to tell you a one-line criteria change was wrong, or a set that reports 95 percent while the agent in production ignores the decision entirely.

Building one: the routing set

The worked example is the decision that gets made most often in our agents, the same kind of per-turn call the coach relies on: given a task brief, which specialist should handle it, or should the parent handle it itself. One wrong answer and the user gets the wrong tool for the job.

One JSON object per line, with a schema stated in a README beside the files:

FieldMeaning
idUnique, prefixed by where it came from
briefThe task as the parent would hand it over, or the user's own words for seeds
targetThe right answer: a specialist, root (the parent handles it) or clarify (ask)
acceptOther targets that are also defensible. Strict scoring uses target only; lenient also accepts these
kindclear, boundary (a hard neighbour pair), vague, root, seed
pairOfFor a vague brief, the clear brief it degrades
sourceFor seeds, the test or eval case the brief came from

The accept field is the one people skip and should not. Many routing questions have two defensible answers. A strict score that calls one of them wrong punishes the judge for a disagreement the humans also have. Reporting strict and lenient side by side tells you how much of the gap is real.

Four files, with their optimism labelled. This is the part I would copy before anything else.

FileCasesHow it was madeHow much to trust it
Development27470 seeds mined from existing tests and eval cases, plus 220 written by hand to cover every specialist, every neighbour boundary, the themes from the agent's own failure log, root-stay cases and paired vague briefsThe criteria were tuned on this file, so its score is optimistic
Holdout74Written by hand after a first read of the errorsNot clean either, and the README says so
Holdout 2124Generated by a model from each specialist's tool list only, never from its descriptionThe fair number
Ambiguous20Context-free briefs like "Fix it." and "Send it."Never scored as a target; only the judge's confidence is read

The second holdout is the trick. The thing being tuned is the specialists' descriptions. A human writing holdout cases has read those descriptions and writes briefs that fit them. A model given only the tool lists has not, so its briefs test whether the descriptions describe the tools, which is the actual question.

What it reported on the development set, 218 delegated cases:

Right specialist, percent of cases
  1. Development, strict90.4%
  2. Development, lenient95.4%
  3. Holdout 2, strict94%
  4. Confidence 0.9 or more96.9%

The last row is the useful one for production. When the judge was at least 90 percent confident, it was right 96.9 percent of the time, across 162 cases. That is a threshold you can route on: below it, fall back to the parent model's own judgment. On the ambiguous set, 9 of 20 briefs came back under 0.6 confidence, which is what you want from a brief that says "Fix it." The weakest slices were two neighbour boundaries, where one kind of request kept landing on the specialist next door. The set names them; the confusion matrix is in the output. That is what "which decision broke" looks like.

And the baseline it was measured against. The same briefs were given to the parent model with its real tool list, recording which tool it would call without running it. It delegated when it should 81.7 percent of the time, picked the right specialist 97.8 percent of the time once it delegated, and landed at 82.6 percent end to end. So the judge was better at deciding whether to delegate, and the model was better at picking whom. The study that produced this set concluded the judge-as-router pattern was not worth adopting. The set is what was left worth keeping.

The rules the set lives under

Golden sets from history

The routing set was written. The other kind is exported.

When a judgment has a human-decided history, the golden set comes from the outcomes, not from the old prompt text. For an answer-selection capability, the set is 200 closed questions with the selections humans finally made, read from the production database through a role that can only SELECT, by a human running one script, because no agent has production access. The exit bar for the new judge was agreement of 90 percent or better with the human-final selections, and the plan noted the risk plainly: the judge is only as good as the golden set, so sample across sources and across answer counts.

The export redacts as the last step before the write, so nothing can be added after it. Email addresses and links become tokens. Author identifiers are hashed with a salt per export, so two exported files cannot be joined back into one identity. Job titles are reduced to a role token. The answer text itself stays whole, because it is published content and the judge has to read it to be worth judging. Each of those rules has a unit test.

A third flavour is the byte-for-byte golden: fifteen email templates, each with its rendered HTML and text saved under two time zones. A 250-line refactor across all fifteen landed as a diff of zero against the goldens. Without them it would have been out of scope for the lane that did it.

Auditing a live judgment

Once a judgment is in production, the golden set stops being the only evidence. We audit against real rows, with a recipe that has held up across 34 live judgments:

  1. Pull up to 50 real rows the judgment actually scored.
  2. Rescore them with the current criteria wording, pulled from source, not from memory.
  3. Hand-review every disagreement.
  4. Try wording variants on the same rows. Zero regressions allowed. Leave the thresholds alone.

That last instruction came from the finding that surprised me most. In every underperforming judgment, the errors sat in a flat band between 0.35 and 0.65 confidence. Naming the observed failure shape in the criteria fixed them. Moving the threshold never did. A threshold tunes how often you are wrong at the margin; the criteria decide where the margin is.

One trap for anyone batching a judge: a batch call that hits a rate limit can return short and silently misalign every later row with the wrong input. Chunk small and assert that the output count matches, or score one row per call.

What the eval suite taught us that sets cannot

The live suite is where the expensive lessons were, and they are worth the two dollars.

Provenance classes. Every case in our suite is labelled either incident-derived, meaning we once got this wrong and no longer do, or rule-derived, meaning we said we would not and still do not. A green incident-derived case is regression evidence. A green rule-derived case proves only our stated intent; it cannot tell you the failure was ever possible. Cases invented from what the rules say are worth much less than cases captured from what actually went wrong, and the label sits at the definition site so nobody counts them the same.

Contamination. The harness ran every case in one invocation under one conversation id, and the agent bootstraps a conversation's history into each new session. So case N saw cases 1 to N−1. One case failed in the full run and passed alone, a false red: the agent was obeying a rule not to re-run a search whose results were already in the thread. Another passed in the full run and failed alone, a false green over the suite's one known live defect. Every full-suite number before that day had been measured under contamination. One case per invocation, a fresh conversation per run, and the spread across three identical runs went from about six points to exactly zero on every case.

Mutation testing against the rule. For two rule-derived cases we deleted the rule they quoted, then inverted it to command the opposite, and ran the case after each mutation. Both stayed green through all four. The behaviour came from the base model, and our instruction layer could not move it in either direction. The cases were kept and relabelled as model-swap guards. The doctrine generalises: a rule-derived case has to be mutation-tested against the rule it quotes, or you cannot tell whether you are measuring your instructions or your model vendor.

The instrument before the measurement. Three independent defects, found within a day of anyone first running the suite, each made a gate report coverage it did not have. The question for every new case became: has this assertion ever been observed to fail for the right reason? And for every new helper: can this matcher fail at all, on an agent that produced no output?

When not to

For a product whose shape was changing weekly, we stopped writing and extending eval suites for a while. The owner's words: we are too early for evals, we do not even know what the product is. What survived that pause was the cheap layer: golden sets of a few dozen cases, and throwaway judge batches of eight or ten examples to pick a threshold. Suites come back when the product settles. Sets never left, because they cost nothing to keep and cents to run.

How to set one up

  1. Pick one decision. A label, a score or a yes/no over text. If it needs the whole agent to evaluate, it is a suite case, not a set case.
  2. Write the schema down, including an accept field for defensible alternatives, and score strict and lenient.
  3. Mine seeds from what already exists, then write cases for every class, every neighbour boundary and a vague twin of each clear brief.
  4. Make a holdout you cannot have been influenced by. Generate it from the inputs the criteria are supposed to describe, not from the criteria.
  5. Label the optimism of every file in the README.
  6. Enforce no real data with a test, and keep the set in sync with the code with the same test.
  7. Run on demand, record the baseline and its date, and compare against the model you are replacing.
  8. Once live, audit on real rows. Fix the criteria, not the threshold.

If you keep golden sets a different way, or have a better answer to the holdout problem than "have a model write it from the tool list", I want to hear it.