Writing

Decision fatigue is real

Every clarifying question an agent asks costs the user a decision. How we moved our chat agent to acting on its best guesses, showing them as editable assumptions on the result, and asking only when a wrong guess is costly.

An agent that asks before it acts feels careful. To the person using it, every question is a decision they have to make before they get anything back. Ask three questions and you have spent three decisions on someone who wanted one answer. Decision fatigue is real, and a polite agent can cause plenty of it.

We spent a day redesigning how our chat agent asks. The short version: it guesses, does the work, shows the guesses on the result where they can be edited, and asks first only when a wrong guess would cost something real.

0/5
Clarifying cards on a broad ask
2/5 before, same prompt, 5 runs each
5/5
Searched on the first turn
0/5 before
18/18
Offers of extra work caught
0/18 before, no false refusals
$0.00003
Cost of the gate per question
one typed judge call, about 250 ms

The problem

The agent had a question tool that could stack up to four steps in one card, a stepper the person walked through before anything happened. Its description said "ask only when blocked". Nothing made it bring a best guess, and nothing limited how often it asked.

So a request like "find me podcasts to pitch" got a card asking what topic, what audience, what tone. Every answer was already guessable from the person's profile and the thread. The person paid for the agent's uncertainty with their attention.

The first fix was not enough

The obvious improvement is to work out an answer and ask the person to confirm it. One click instead of typing. When several things need confirming, list them together and make each row editable in place, pre-filled with the guess, with one submit.

That is better, and we kept the component. But it still puts a decision in front of the person before they see any value. A confirm step is a question with a default answer. By the end of the same day we had moved past it.

The approach

1. Act, then show the assumptions on the output

The agent works out what it needs from memory, the account and the thread, and does the work at once. The result carries an Assumptions strip: one row per guess, each with its value and where it came from.

Getting value costs zero decisions. A wrong guess costs one edit, made where its effect is visible. We measured the edit loop on real drafts:

Step after the person edits a rowTime
Saved to memory52 ms
Draft body rewritten4.8 s
Strip unlocked, 8 runs5.5 to 7.0 s

2. Ask first only when a wrong guess is costly

A blocking question is allowed in three cases:

Any other question is refused, and the agent is told to act on its best guess and show it in the assumptions.

3. One component, two placements

The blocking question is the same assumptions list, shown before the work instead of on it. All rows at once, no stepper. Guessed rows are pre-filled and editable. A row with no basis for a guess is an empty field, which is how open questions survive: as one empty row, not a separate mechanism. At most three rows, one submit.

4. Hard limits, enforced in code

Rules in a tool description are requests. We put the limits in a gate that runs inside the question tool, before the card is shown:

The refusal goes back to the model as the tool result, so it reads like a one-line nudge from a coach rather than another rule in the prompt. The judge returns typed booleans with a probability, not prose to parse.

5. Learn from every edit

An edited assumption is saved to memory, so the next run's guess is right by default. Over time, blocking questions should trend toward zero.

What we measured

We built a small labelled set of real moments, each marked "act and show", "ask first" or "already known", and kept a held-out set we never tuned on.

SetCasesAsk precisionMedian latencyCost per check
Tuning321.00246 ms$0.000028
Holdout100.75237 ms$0.000028

Then a live probe, five runs per arm on the broad podcast ask:

MetricBeforeAfter
Clarifying cards shown2/50/5
Searched on turn 10/55/5

What broke along the way

The model asked in prose instead. With the gate alone, cards dropped to 1 in 5, but in 3 of 5 runs the model asked the same question as plain text and did not search. Blocking the tool does not block the habit. One instruction change fixed it: do the broadest sensible version of the request and state the assumption, never ask in prose. Turn-1 searches went to 5/5.

Offers looked like guessable questions. A card offering extra work ("want me to draft pitches for these?") is guessable too, so the gate told the model to act on its best guess, and it did the extra work unasked. A question about how many items qualified ended with two drafted emails nobody requested. The fix was one more boolean in the same judge call: is this an offer beyond the person's latest message? Offers are refused with an instruction to mention them in one line as an optional next step. Offers caught went from 0 of 18 to 18 of 18, with no ordinary questions misread as offers.

Free-text reasons read like options. Each guessed row first carried a sentence from the model explaining the guess. It sat between the label and the field, ran long, and read like another choice. We replaced it with a closed set of sources shown as a small tag ("from your profile", "from this thread"), plus an optional one-sentence reason behind a "Why?" button, only for judgement calls.

If you run an agent that asks questions

If you have measured how often your agent asks before acting, I would like to see your numbers.