Decision fatigue is real
Every clarifying question an agent asks costs the user a decision. How we moved our chat agent to acting on its best guesses, showing them as editable assumptions on the result, and asking only when a wrong guess is costly.
An agent that asks before it acts feels careful. To the person using it, every question is a decision they have to make before they get anything back. Ask three questions and you have spent three decisions on someone who wanted one answer. Decision fatigue is real, and a polite agent can cause plenty of it.
We spent a day redesigning how our chat agent asks. The short version: it guesses, does the work, shows the guesses on the result where they can be edited, and asks first only when a wrong guess would cost something real.
The problem
The agent had a question tool that could stack up to four steps in one card, a stepper the person walked through before anything happened. Its description said "ask only when blocked". Nothing made it bring a best guess, and nothing limited how often it asked.
So a request like "find me podcasts to pitch" got a card asking what topic, what audience, what tone. Every answer was already guessable from the person's profile and the thread. The person paid for the agent's uncertainty with their attention.
The first fix was not enough
The obvious improvement is to work out an answer and ask the person to confirm it. One click instead of typing. When several things need confirming, list them together and make each row editable in place, pre-filled with the guess, with one submit.
That is better, and we kept the component. But it still puts a decision in front of the person before they see any value. A confirm step is a question with a default answer. By the end of the same day we had moved past it.
The approach
1. Act, then show the assumptions on the output
The agent works out what it needs from memory, the account and the thread, and does the work at once. The result carries an Assumptions strip: one row per guess, each with its value and where it came from.
Getting value costs zero decisions. A wrong guess costs one edit, made where its effect is visible. We measured the edit loop on real drafts:
| Step after the person edits a row | Time |
|---|---|
| Saved to memory | 52 ms |
| Draft body rewritten | 4.8 s |
| Strip unlocked, 8 runs | 5.5 to 7.0 s |
2. Ask first only when a wrong guess is costly
A blocking question is allowed in three cases:
- Irreversible or spending: sending, buying, publishing, scheduling.
- Unguessable: a fact only the person knows, such as an embargo date or a number.
- Low confidence on something load-bearing: the whole output hinges on it.
Any other question is refused, and the agent is told to act on its best guess and show it in the assumptions.
3. One component, two placements
The blocking question is the same assumptions list, shown before the work instead of on it. All rows at once, no stepper. Guessed rows are pre-filled and editable. A row with no basis for a guess is an empty field, which is how open questions survive: as one empty row, not a separate mechanism. At most three rows, one submit.
4. Hard limits, enforced in code
Rules in a tool description are requests. We put the limits in a gate that runs inside the question tool, before the card is shown:
- One blocking question per task. A per-turn counter refuses the second.
- Never ask what is already known. One typed judge call reads the recent thread, the saved memories and the card, and refuses anything the context answers.
- Silence means proceed. A skipped question, a new message or no answer means the agent goes ahead on its guesses and says so in one line.
The refusal goes back to the model as the tool result, so it reads like a one-line nudge from a coach rather than another rule in the prompt. The judge returns typed booleans with a probability, not prose to parse.
5. Learn from every edit
An edited assumption is saved to memory, so the next run's guess is right by default. Over time, blocking questions should trend toward zero.
What we measured
We built a small labelled set of real moments, each marked "act and show", "ask first" or "already known", and kept a held-out set we never tuned on.
| Set | Cases | Ask precision | Median latency | Cost per check |
|---|---|---|---|---|
| Tuning | 32 | 1.00 | 246 ms | $0.000028 |
| Holdout | 10 | 0.75 | 237 ms | $0.000028 |
Then a live probe, five runs per arm on the broad podcast ask:
| Metric | Before | After |
|---|---|---|
| Clarifying cards shown | 2/5 | 0/5 |
| Searched on turn 1 | 0/5 | 5/5 |
What broke along the way
The model asked in prose instead. With the gate alone, cards dropped to 1 in 5, but in 3 of 5 runs the model asked the same question as plain text and did not search. Blocking the tool does not block the habit. One instruction change fixed it: do the broadest sensible version of the request and state the assumption, never ask in prose. Turn-1 searches went to 5/5.
Offers looked like guessable questions. A card offering extra work ("want me to draft pitches for these?") is guessable too, so the gate told the model to act on its best guess, and it did the extra work unasked. A question about how many items qualified ended with two drafted emails nobody requested. The fix was one more boolean in the same judge call: is this an offer beyond the person's latest message? Offers are refused with an instruction to mention them in one line as an optional next step. Offers caught went from 0 of 18 to 18 of 18, with no ordinary questions misread as offers.
Free-text reasons read like options. Each guessed row first carried a sentence from the model explaining the guess. It sat between the label and the field, ran long, and read like another choice. We replaced it with a closed set of sources shown as a small tag ("from your profile", "from this thread"), plus an optional one-sentence reason behind a "Why?" button, only for judgement calls.
If you run an agent that asks questions
- Count the decisions your agent costs a person before it delivers anything. Each question is one.
- Act first and show the guesses on the result, editable in place. Make the edit rework the output.
- Allow a blocking question only when a wrong guess is costly: irreversible or spending steps, facts only the person knows, or low confidence on something the whole output depends on.
- Enforce the limits in code, inside the tool: one question per task, never ask what is known, silence means proceed.
- Watch for the workarounds: questions moving into prose, and offers of extra work slipping through as guesses.
- Save every correction so the same guess is right next time.
If you have measured how often your agent asks before acting, I would like to see your numbers.