Stop asking a language model yes-or-no questions
A typed decision model, Jev, answers labels, scores and booleans with a probability for a fraction of a cent. How it unlocked flows we could not afford before, where it does not fit, and the method we use to find the next gain.
Most of what we used to ask a language model was not language. Is this answer relevant. Which of eight kinds is this support message. Does this profile belong to a journalist. Does this draft break rule 14. We asked a text model, parsed "YES" off the front of its reply, and paid for a paragraph of reasoning nobody read.
In September we moved those decisions onto Jev, a model that takes a state and a set of typed questions, a choice over named options, a score on an ordered rubric, or a boolean, and returns an answer to each with a probability. It writes no prose, calls no tools, and extracts nothing. It is reached through the same AI Gateway as everything else. It has changed how we build more than any other single tool this year, and this post is about why, what it opened up, and how we go about finding the next place it belongs.
Why it mattered
The price is the headline, and it is real: Vercel's own numbers put it at up to 190 times faster and 440 times cheaper than a text model on the same decisions, and ours are in that range. But price is not what changed the design. Three properties did.
A probability you can route on. A text model asked for a yes gives you a yes. Jev gives you 0.93, or 0.51. That turns every decision into a threshold you can tune and a fallback you can trigger. Confident, act. Unsure, send it to the expensive model, or to a person. Our rule of thumb from several measured sites is that the errors live in a flat band between 0.35 and 0.65, and the fix is almost always to name the failure shape in the criteria rather than to move the threshold.
No prose means nothing to hallucinate. A whole class of defensive code disappeared: the quote-verification helpers that existed to catch invented evidence, the "did the model actually run" guards keyed on token counts, the startsWith("YES") parsers. If the output is a label from a set you defined, the model cannot make up a field you did not ask for.
Cheap enough to put in front of everything. When a decision costs a few hundredths of a cent and a third of a second, you stop rationing it. A "no" from a Jev gate skips the language-model call it fronts, so the gate pays for itself on the first skipped call. That one pattern, decide first and generate only when the decision says so, is behind most of the flows below.
What it unlocked
These are flows that either did not exist or were too expensive to run everywhere before.
- The coach. One line of per-turn guidance for the agent, chosen by a Jev classification of the incoming message, costing nothing on turns where no rule applies. I wrote that one up separately.
- Identity gates. A search by a person's name returns namesakes, and a namesake's interview was ending up under the wrong expert. The old defence was a prompt line telling the extraction model to "confirm identity". Now every finding and every discovered transcript is judged "about this person" before anything is mined or stored, and a skipped one logs its probability.
- Relevance gates in front of extraction. A query that is not about the thing at all used to get the full extraction call anyway. Now one boolean runs first; irrelevant means an empty result and no language-model call.
- Guardrails as one boolean per rule. An answer that is derived from a knowledge base is checked against each content rule separately, fail-closed. The rule ids that answered true are the finding. The same shape runs 93 content rules against a draft article: the first check costs $0.0012, each check after an edit about $0.0004, because paragraphs are cached and only the author-visible rules run live.
- A second opinion on a judgment with tools. Moderation needs a tool loop to gather evidence, so the decision stays on a text model. Jev reads the moderator's own output and answers two booleans; a disagreement becomes a manual review either way. Cheap insurance on the one judgment that publishes to a list.
- Cheap-first generation. For generation tasks the model cannot do, it still gates: run a cheap model, have Jev check the output against a few criteria, retry twice, fall back to the expensive model. Measured on 40 inputs with a blind judge, one cheap model matched production's quality at 4.9 times lower cost and fell back once in forty. Three other cheap models scored far worse, and the gate accepted most of their poor output. The lesson in the decision record: a stricter gate does not make a weak model safe; pick the model by measurement, not price.
- Finding that production was over-editing. An "does this need an edit" gate in front of a rewriting step returned 63 percent of answers untouched, and the blind judge scored the result higher than production. The expensive step had been changing quotes that were fine.
- Support routing. One choice over eight kinds of support message, and one choice over four specialist desks, replacing two text-model calls with a probability each, scored against the corrections staff make.
And one from this site: the post before this one classified 218 of my own messages by tone, concreteness and whether they gave a reason. Four questions, 218 calls, a few cents, two minutes.
How we find the next one
This is the part I would hand to another team, because the gains are not where you would first look.
1. Inventory every model call, one row each. Three read-through passes over the code, one row per call site, with columns: what typed output does it produce, what prose output, does it use tools, is there already a comparer scoring it, and the fit. The first pass over one codebase found about 110 sites. The second, a week later, found almost all the remainder were generation, extraction, tool loops or embeddings.
2. Ask the right question. The first inventory asked "does the prompt generate?" and missed five sites. The second asked "does the consumer use anything but the decision?" and found them: a spam filter whose consumer only read an id list, a relevance filter returning indices, a profile-match validator returning a distance, a publication filter returning URLs, a category assigner returning a top five. Each produced prose the consumer threw away.
3. Classify the fit, not just yes or no. Five verdicts cover everything we have seen:
| Verdict | When | Example |
|---|---|---|
| Switch | Typed output, no tools, nothing a person reads | A PASS/FLAG review with a confidence |
| Split | One call did a decision and an extraction | Decide "required?" first; extract the requirement text only when true |
| Partial | The typed part is small beside an extraction the model must do anyway | Three booleans riding on a demographics extraction: not worth the second call |
| Guardrail | The decision needs a tool loop | A second pass over the moderator's output |
| No | Prose a person reads, or pure extraction | A one-sentence "why" rendered on a row |
4. Know what it cannot do, and do not try. Anything a customer reads as a sentence. Extraction of a span from the input. A judgment whose evidence comes from a tool loop. A deterministic gate the design already took the model out of. And a vendor's calibrated score for a specific job, such as whether text reads as machine-written: Jev is not a detector and should not be asked to be one.
5. Measure the way the decision deserves. Smoke probes, one real input per shape against the live gateway, when the ask is to adopt and measure in production. Comparisons with a blind judge when the ask is a model choice. For cheap, reversible judgments with a comparer already in place, the owner's rule was blunt: just switch, we will test in production, all this extra infrastructure is not needed. Shadow windows were kept for the one judgment that sits behind a parity contract.
6. Audit once live. Fifty real rows, rescored with the current wording from source, hand-reviewed, with variants tried on the same rows, zero regressions allowed and the thresholds left alone. Across 34 live judgments that recipe has held.
7. Record every judgment. Each writes one decision row naming its model and the hash of its question set, so a change in behaviour can be traced to a change in criteria.
What did not work
Honesty about the misses is part of the method.
- Routing an agent's subagents by Jev. The study was careful, with a dev set and two holdouts. Jev was better than the parent model at deciding whether to delegate and worse at picking whom. The conclusion was that the router pattern was not worth adopting; the golden set it produced was the part worth keeping.
- Per-input routing for generation. Trying to predict from Jev features which inputs a cheap model would handle well did not work: the same input passed and failed across runs.
- Two first-draft gates rejected production's own outputs. Calibrate a gate on what production already produces before you let it block anything.
- One comparator is nondeterministic. A semantic-redundancy judgment changed its answer on 15 of 50 rows across runs. That needs a structural fix, not better wording.
- Operational traps. One gateway call stalled for 302 seconds, so every call now aborts at 20 seconds and retries. A batch that hits a rate limit can return short and silently misalign later rows, so batches are small and the output count is asserted.
Where it leaves us
Decisions used to be expensive, so we made few of them and made them in prose. Now a decision costs less than the logging of it, so we make them everywhere: in front of generation, after generation, per turn in the agent, per rule in a draft, per candidate in a search. The language model does what only it can do, which is write and read freely. Everything that was secretly a label is a label again.
If you take one thing: list every place you ask a model for a yes, a score or a pick, and look at what the caller does with the words that come back. In our code, most of the time, it threw them away.