Writing

Stop asking a language model yes-or-no questions

A typed decision model, Jev, answers labels, scores and booleans with a probability for a fraction of a cent. How it unlocked flows we could not afford before, where it does not fit, and the method we use to find the next gain.

Most of what we used to ask a language model was not language. Is this answer relevant. Which of eight kinds is this support message. Does this profile belong to a journalist. Does this draft break rule 14. We asked a text model, parsed "YES" off the front of its reply, and paid for a paragraph of reasoning nobody read.

In September we moved those decisions onto Jev, a model that takes a state and a set of typed questions, a choice over named options, a score on an ordered rubric, or a boolean, and returns an answer to each with a probability. It writes no prose, calls no tools, and extracts nothing. It is reached through the same AI Gateway as everything else. It has changed how we build more than any other single tool this year, and this post is about why, what it opened up, and how we go about finding the next place it belongs.

21 → 34
Decision sites moved in the first app
over about a week, no dual path kept
~$0.00004
Cost per judgment
850 to 1,000 input tokens at $0.042 per million; output is free
200–730 ms
Latency per judgment
several questions answered in one call
2 vs 32
Wrongful removals on one eligibility decision
Jev against the text model it replaced, 852 answers

Why it mattered

The price is the headline, and it is real: Vercel's own numbers put it at up to 190 times faster and 440 times cheaper than a text model on the same decisions, and ours are in that range. But price is not what changed the design. Three properties did.

A probability you can route on. A text model asked for a yes gives you a yes. Jev gives you 0.93, or 0.51. That turns every decision into a threshold you can tune and a fallback you can trigger. Confident, act. Unsure, send it to the expensive model, or to a person. Our rule of thumb from several measured sites is that the errors live in a flat band between 0.35 and 0.65, and the fix is almost always to name the failure shape in the criteria rather than to move the threshold.

No prose means nothing to hallucinate. A whole class of defensive code disappeared: the quote-verification helpers that existed to catch invented evidence, the "did the model actually run" guards keyed on token counts, the startsWith("YES") parsers. If the output is a label from a set you defined, the model cannot make up a field you did not ask for.

Cheap enough to put in front of everything. When a decision costs a few hundredths of a cent and a third of a second, you stop rationing it. A "no" from a Jev gate skips the language-model call it fronts, so the gate pays for itself on the first skipped call. That one pattern, decide first and generate only when the decision says so, is behind most of the flows below.

What it unlocked

These are flows that either did not exist or were too expensive to run everywhere before.

And one from this site: the post before this one classified 218 of my own messages by tone, concreteness and whether they gave a reason. Four questions, 218 calls, a few cents, two minutes.

How we find the next one

This is the part I would hand to another team, because the gains are not where you would first look.

1. Inventory every model call, one row each. Three read-through passes over the code, one row per call site, with columns: what typed output does it produce, what prose output, does it use tools, is there already a comparer scoring it, and the fit. The first pass over one codebase found about 110 sites. The second, a week later, found almost all the remainder were generation, extraction, tool loops or embeddings.

2. Ask the right question. The first inventory asked "does the prompt generate?" and missed five sites. The second asked "does the consumer use anything but the decision?" and found them: a spam filter whose consumer only read an id list, a relevance filter returning indices, a profile-match validator returning a distance, a publication filter returning URLs, a category assigner returning a top five. Each produced prose the consumer threw away.

3. Classify the fit, not just yes or no. Five verdicts cover everything we have seen:

VerdictWhenExample
SwitchTyped output, no tools, nothing a person readsA PASS/FLAG review with a confidence
SplitOne call did a decision and an extractionDecide "required?" first; extract the requirement text only when true
PartialThe typed part is small beside an extraction the model must do anywayThree booleans riding on a demographics extraction: not worth the second call
GuardrailThe decision needs a tool loopA second pass over the moderator's output
NoProse a person reads, or pure extractionA one-sentence "why" rendered on a row

4. Know what it cannot do, and do not try. Anything a customer reads as a sentence. Extraction of a span from the input. A judgment whose evidence comes from a tool loop. A deterministic gate the design already took the model out of. And a vendor's calibrated score for a specific job, such as whether text reads as machine-written: Jev is not a detector and should not be asked to be one.

5. Measure the way the decision deserves. Smoke probes, one real input per shape against the live gateway, when the ask is to adopt and measure in production. Comparisons with a blind judge when the ask is a model choice. For cheap, reversible judgments with a comparer already in place, the owner's rule was blunt: just switch, we will test in production, all this extra infrastructure is not needed. Shadow windows were kept for the one judgment that sits behind a parity contract.

6. Audit once live. Fifty real rows, rescored with the current wording from source, hand-reviewed, with variants tried on the same rows, zero regressions allowed and the thresholds left alone. Across 34 live judgments that recipe has held.

7. Record every judgment. Each writes one decision row naming its model and the hash of its question set, so a change in behaviour can be traced to a change in criteria.

What did not work

Honesty about the misses is part of the method.

Where it leaves us

Decisions used to be expensive, so we made few of them and made them in prose. Now a decision costs less than the logging of it, so we make them everywhere: in front of generation, after generation, per turn in the agent, per rule in a draft, per candidate in a search. The language model does what only it can do, which is write and read freely. Everything that was secretly a label is a label again.

If you take one thing: list every place you ask a model for a yes, a score or a pick, and look at what the caller does with the words that come back. In our code, most of the time, it threw them away.