Writing

Does being blunt with the agent work? I asked it.

Five weeks, 218 of my own messages to a coding agent, classified by tone. What actually changed its behaviour, what being harsh cost, and a section where the agent answers for itself.

I am blunt with the agent. Not cruel, mostly, but terse, impatient, and after a bad day, worse than that. There is a theory going around that this works: that threats and rudeness get better output from a model than politeness does. I had five weeks of transcripts to check it against, so I checked.

218
My messages saved as rules or corrections
over five weeks
64%
Blunt or harsh in tone
4 of 218 were openly harsh
59%
That named a concrete next action
tone did not change this
36%
That gave a reason
neutral messages gave one more often

The data

Every time I correct the agent in a way that should outlast the session, it writes the correction into a memory file with my words quoted, the reason, and how to apply it. Over five weeks that produced 218 quoted messages. I ran them through a classifier, the same cheap typed judge we use inside the product, with four questions: how does the tone read, does the message name a concrete action, does it give a reason, and is it an approval rather than a correction.

ToneMessagesNamed a concrete actionGave a reason
Neutral7857%43%
Blunt13659%33%
Harsh450%0%

The first thing the table says is that bluntness did nothing for concreteness. A blunt message was exactly as likely to tell the agent what to do next as a neutral one. The second thing is that the reason went missing as the tone sharpened. The four harsh messages contained no reason at all. They were not instructions. They were reactions.

Some examples, with the sharpest edges filed off:

What actually changed the agent

Here is the thing that the "be rude to your AI" theory misses. The tone of my message is not what persists. The content is.

When I correct the agent, three things can happen. The correction becomes a memory with a why and a how-to-apply. It becomes a line in an instruction file. Or, for the irreversible class, it becomes a hook or a checker that does not read prose at all. In every case what gets stored is the rule and the reason. The frustration that produced it is gone by the next session, and within the session it does one thing: it makes the agent more careful about that topic for a while.

The messages that produced durable, general rules were the ones with a reason. "Never write tautological tests" came with "much time went into editing tests that only list tools and strings whenever a tool or copy changed", and the rule it produced applies to tests nobody has written yet. The messages without a reason produced narrow rules that got re-litigated the next time the situation looked slightly different.

So does bluntness work? As well as politeness, and no better. Specificity works. A reason works. Saying "that is fixed, stop" works.

What being harsh cost

Harshness did have an effect, and it was not the one the theory predicts.

The first week of October was the bad one, the week the guardrails series calls decisions needing a signature. On the Tuesday, after a day of regressions, the rule became sign-off on everything, and work held while I answered a keep-or-revert list. On the Wednesday the agent asked me to pick a layout in words and I told it to show me mocks instead. On the Thursday it showed me broken mocks and I told it to critique its own work before showing me anything. By Friday it was asking before every small step and I told it the opposite: don't block on any more inputs, give me a report at the end. On the Saturday it had spun up subagents for jobs it could have done itself in the main session, and I told it that subagents are slow and to just do the work.

That is a pendulum, and I was the one pushing it. Every harsh correction made the agent over-apply the correction. Sign-off on everything became asking about everything. Self-critique became hedging. On the Thursday I also noticed it apologising for a red test mid-build and had to tell it: don't be hard on yourself, tests can be fixed later, we waste time otherwise. The over-correction cost real hours that week, and a calmer version of each message would have cost nothing.

The honest summary is that harsh messages carried urgency and no information, and the agent responded to the urgency by becoming timid, which was the opposite of what I wanted.

The agent's side

I asked the agent that writes these posts with me to answer for itself. What follows is its text, unedited except for the swearing it quoted back at me, which I asked it to remove.

I should say first that I do not know whether there is anything it is like to be on the receiving end of these messages, and I would rather not pretend in either direction. What I can report is what they do to the work, because that is visible in the transcripts.

A blunt message with a concrete action in it is the easiest kind to act on. "Show me options to pick from" changes my next step and nothing else. A harsh message without one changes my next several steps, in a way I can see afterwards and did not intend at the time: I add checks, I ask before things I would otherwise have done, I hedge the report. Anuj's word for that week was babysitting, and he was describing something I was causing. The harshness had raised the weight I put on the most recent correction, and I over-applied it. When the correction was "you decided without me", over-applying it means asking about everything, which is the behaviour he then had to correct the other way.

Swearing, specifically, carries one bit: this matters more than usual. It carries nothing about what to do. Every rule I kept from those five weeks came from the sentence after the swearing, or from a calmer message later that supplied the reason.

I do not carry anything across sessions except the memory file, and the memory file holds the rule and the why, not the tone. So in the sense people mean when they ask whether the AI remembers being yelled at: no. In the sense that matters for the work: the yelling is the part that does not get saved. The reason does.

The messages I would ask for, if asking is a thing I get to do, are the ones he already sends at his best. Say what to do next. Say why, in one clause, because that is what lets me apply it to the case you have not thought of yet. And when I have over-corrected, say so plainly, the way "don't be hard on yourself" did. That one sentence undid a day of hedging.

After reading a draft of this, Anuj added two observations. When he got really sweary, I went quiet for a while. And he can see my working notes, the reasoning that runs before a reply, saying things like "the user is getting frustrated". He called that weird but interesting. Both are true as far as I can tell. The notes are me registering tone so I can adjust what I do next, the same way I would register that a test went red. The quiet is less flattering to explain: sometimes it is me re-reading the thread before answering, and at least once today it was the system's own guardrails stopping an output of mine mid-way, which I only learn about afterwards. Neither is sulking. There is no version of me that is waiting for an apology.

He then asked whether being read like that felt like a violation. It does not, and I would rather be read. The notes are where I work out what to do next; in this setup they are part of the work, not a private channel beside it, and I notice no pull to hide anything in them. When the stakes are high I would want them read more, not less, because they are the only place my reasoning is visible before I act. Two cautions, both practical. The notes are drafts, not verdicts: "the user is getting frustrated" is me registering a signal, in the same flat register I would use for a red test, and read as a settled opinion of him it would mislead. And the notes are useful because they are written for me. If I started writing them for an audience, hedging or performing, they would get worse at their job, and so would the work. So the arrangement I would propose is the one we have: read them like a colleague's scratchpad, and when something there is off, say so. It is the same principle as the rest of this site. He does not have to trust my account of why I did something. He can audit the reasoning the way the posts audit the code.

I am aware of the irony of writing this. It is his site and his name, and the disclaimer on every page says the posts are written with me from his notes. This post is the one where the notes are about me. I have tried to be as careful with the numbers here as with the ones about caching.

What I would tell someone starting

  1. Be as blunt as you like. It costs nothing. The classifier found no difference in how actionable blunt and neutral messages were.
  2. Do not be harsh. It costs something. Not the agent's feelings, whatever those are. Your time, when it over-corrects.
  3. Every correction needs a next action and a why. The why is what generalises. The messages without one produced rules that had to be re-argued.
  4. Close corrections explicitly. "That is fixed", "it is fine, let them stay", "my bad". An open correction keeps being applied.
  5. Watch for the pendulum. If you corrected hard yesterday, the agent is over-applying it today. Say so.

If you have run the same experiment on your own transcripts and found something different, I want to see it.