The agent kept a diary. Nobody read it.
We tried to make a product observe itself from inside its own code. It cost tokens on every turn and changed nothing. A scheduled agent that runs once a week, outside the product, did what the observer could not.
There is an obvious idea once you run an AI agent in production: let it watch itself. Give it a place to write down what went wrong, read that back to it, and it will improve. We did that. It did not work, and the way it failed taught me where observation belongs.
The observer in the code
The agent that runs our chat product had a diary tool. When something went badly in a conversation, a search that returned nothing, a tool it should not have called, a rule it noticed it had broken, it wrote a short note to a database table. On every later turn, the most recent notes were appended to the prompt as a recall block, so the agent would remember its lessons.
It is a tidy design and it has three problems, each of which we found the slow way.
Nobody reads a diary that writes itself. In two weeks the agent wrote 109 entries. No human read them as they arrived, because nothing asked anyone to. When we finally sat down and read all of them, in one afternoon, they clustered into eight fixes. Most of the rest were the same lesson rephrased, a signal that the fix had never landed, not that a new one was needed.
The reader pays every turn. The recall block rendered the last thirty entries as JSON with timestamps: 2,115 tokens, appended to every turn of every session, chat and API, signed in and anonymous. Rewriting it as twelve plain lines cut that to 581. Archiving the entries whose lessons had become policy cut it to zero. The observer had been the single largest always-on cost in the prompt, and the agent it was meant to help had been paying it to remember things nobody had acted on.
Observation inside the product wants to become action inside the product. Once the diary existed, the next proposal was natural: have the agent review its own entries, draft fixes, open pull requests. I said no, and wrote it down as a boundary: the code emits telemetry and nothing else; the review, fix and pull-request loop lives off the platform. A product that rewrites itself from inside its own request path is a product you cannot reason about, and every step of that loop would run at user latency and user cost.
So the diary became a cheap, read-only signal, the kind the coach is built from. The question was where to put the loop that reads it.
The report that works
The answer was a scheduled agent, outside the product, with read-only access to everything, that runs once a week and writes for people.
Concretely: a Claude Code skill, a few hundred lines of instructions, run by a scheduler every Monday morning. It reads the production database through a role that can only SELECT, the payment provider through a restricted key, the hosting platform's observability API, the support tool, and the git history since its last run. It writes a report of 1,000 to 1,400 words in ten fixed sections, renders it to a static page on one hosting project behind authentication, and emails the founders a summary of five to eight lines with a link. It opens small pull requests for instrumentation gaps it found. Then it edits its own instructions with what it got wrong, and records that in a changelog at the bottom of the file.
It is not a product feature. No user ever waits on it. It costs one run a week instead of a slice of every turn. And it has a reader, because it is addressed to someone.
It took six drafts
The first run produced six versions in one day. The first went to all three founders and was rejected in one sentence that I have kept verbatim as a writing rule: it reads very AI written, packed with too much detail; I want takeaways and actionable items, information not just data.
Each later draft produced a rule, and the rules are the interesting part.
- Answer it yourself before asking anyone. Version three asked the engineering team two questions that were answerable from git history and logs. The owner's response: you have access to git history, I should not be the one to track it down, everything you can answer, that is the whole point. There is now a checklist. Find the hour a metric stepped, then look at what deployed that hour. Split an anomaly by cohort before describing it. Compare shares before claiming a cause: errors inside a window as a share of all errors, against the window's share of minutes. Only what survives the checklist may appear as a question, and it has to say what was ruled out.
- A rate before a claim. Version five called a model fallback path broken. It had a 98 percent success rate; the "errors" were SDK warnings and retry payload dumps counted as failures. Nothing is called broken now without the measured rate and the instrument that measured it.
- Know what your instrument measures. Version six said pages had become slower. The metric was function time to first byte, which under the framework's caching model measures the dynamic hole of a page, not the page. The static shell comes from the CDN. The finding was that a page still invoked a function per view for a small check, which is a different and more useful sentence.
- The email is about the business. When the format changed, the summary email started describing its own process: hosting, attachments, what the skill had learned. The owner: the email should not talk so much about this process. Those sentences now go in the changelog and nowhere else.
Corrections per run are counted. The first run had four. The target is zero by the fourth run, and if the count does not fall the pre-send self-check gets stricter, reading the draft as the owner would and asking of every claim: what is the rate, what is the instrument, what is the deploy.
The failure that proved the design
On its second scheduled Monday, the run crashed at startup. The scheduler had launched it with a default limit of 256 open files, which the tool exceeded while loading. Nobody knew for two days, because a report that does not arrive looks exactly like a Monday when you forgot to check.
The owner's reaction was the right one: and I did not even know about it; you can send an email when something goes wrong, no? The job now runs under a wrapper that raises the file limit, caps the run at four hours, writes a marker when a report is actually sent, and emails one person if the run exits non-zero, hangs, or finishes without sending. On the other six days it checks once that the latest Monday has a report and alerts if it does not.
An observer inside the product would never have had this failure, and that is not a point in its favour. It would have failed silently inside a request, as the diary did, and nothing would have noticed the silence.
What the report grew into
Because it is a document with a reader, requests arrive as edits to the instructions, and the instructions compound.
- Prevalence on every issue. Not "some customers", but how many of how many, how often per week, and the money at stake if any. One account affected is still reported; its prevalence says it is one.
- Consequences before confidence. Every proposed product action gets its own analysis: who it touches, the upside and how it was derived, the risks, the guard metric and rollback, and the cheapest way to de-risk it. The analysis can overturn the proposal, and a held proposal stays visible with the reason. Twice it has corrected the report's own claims before they went out.
- Show the fixes, not the count. When the report found help-centre articles contradicting the product, it said so as a number. The owner asked for all of them, with proposed edits and a button. The page now renders each edit, before struck through and after in green, with the code evidence, and one button a founder presses to apply the checked ones. The run never presses it.
- A cleanup sweep. Each week it also opens one pull request per repository removing comments that contradict the code beside them and tests that only restate the code, scoped to the ripple of that week's merges into files nobody touched. That one is a post of its own.
Where observation belongs
The diary taught the agent to remember. It did not teach anyone anything, because remembering is not the same as someone reading. The weekly run has a worse memory, one state file and a notes file, and a far better record, because every one of its findings has been read by a person who could act on it, and every one of its mistakes is written into the instructions for the next run.
If you are tempted to make your product observe itself, ask who reads the observation, what it costs the user per turn, and what happens when it is silent. Then consider putting the observer outside, on a schedule, with read-only credentials and a human at the other end of an email.