Writing

Have an agent use the app and say what hurts

Turning "let an agent test the app and report friction" from a one-off into a system. Five parts, a ledger that remembers, deterministic probes before any model, and the dogfood rounds that feed it.

The request sounded simple: have an agent use the app the way a person would and report the friction. We had done it once, by hand, with a script that rendered every email template and an agent that looked at them. It found things. It also had no rubric, no memory of what it found last time, and no way to tell run two from run one. The second time the same request came back with a different noun, screens instead of emails, we made it a system.

5
Parts of a sweep
surface, harness, probes, rubric, ledger
52
Templates in the first sweep
rendered at two widths, checked by probes and a rubric
4
Fixtures per list screen
empty, few, many, long values
4–6
False positives in the first dogfood round
now caught by a verification round and brief rules

The shape

A sweep is five parts. An instance names all five; shared plumbing makes an instance a few files.

surface  →  harness  →  probes  →  rubric  →  ledger

The rules that make it a system

A sweep with no ledger is a demo. If run two cannot say what changed since run one, it is not done. The ledger is tracked in git, so a pull request that fixes a finding shows the finding leaving the file, and a reviewer can see it.

Probes first. If a deterministic check can answer it, do not ask a model. The judge is for what the DOM cannot say. A person or a vision model is for what the judge cannot see.

Only units the run swept can close a finding. A partial run, restricted to one screen, is therefore safe. It cannot accidentally mark the rest fixed.

Accepted findings stay in the ledger with a reason, so a nit is not re-litigated every run. They close like any other when they disappear.

Fixtures over mocks. The harness renders from props. There is no if (MOCK) in a data loader anywhere. That rule turned out to be the one with the largest consequences, and it gets its own post.

Changing a rubric question is a pull request. So when a score moves, the product moved.

Three sweeps

Emails. Every template rendered with its preview props, checked as markup, opened in a browser at 720 and 375 pixels for overflow and small targets, run through the accessibility checker, then the email rubric: how many calls to action, whether their labels are generic, whether the subject matches the body, whether the opening says why the reader got it. The first run found the real root causes behind a dozen findings: a link colour with insufficient contrast in the shared template, and a box-sizing rule in the body. Fixed once, closed everywhere.

Screens. Every registered screen times every fixture times two viewports, drawn by the preview harness inside the real application frame. Probes measure the screen's own content, not the frame around it, so a navigation defect is not reported once per screen; the frame is registered as its own unit and swept once. The screen rubric asks about competing primary actions, unexplained internal shorthand, placeholder copy, and whether an empty state says why it is empty and what to do next. Adding a screen to the registry adds it to the sweep.

Flows. The one an agent drives by hand. Personas and tasks are the surface, with tasks worded as the persona would think them, never as routes. The agent signs in as the persona on a preview deployment, attempts each task using only what the persona would know, and journals every moment of friction with a type, hesitation, dead end, surprise, unclear result, unreachable, and a key naming the thing rather than the moment. The journal is then triaged through a friction rubric that answers two questions, who would hit it and what it costs, and reconciled into the same ledger. Tester-only entries are dropped. Meant to run weekly as a scheduled agent.

What the dogfood rounds taught

Before the sweeps there were dogfood rounds: five functional agents in parallel, each on one surface, plus a latency agent, plus a human pass with real prompts. The process document that came out of round one is mostly a list of what went wrong, which is the useful kind.

What it does not do

No golden-screenshot pixel diffing: the ledger is about findings, not pixels, and screenshots are artefacts for a reviewer. No vision model in the loop: when one is wanted, it is a generator in front of the rubric, writing the same journal the flows sweep uses, and triaged the same way. No component storybook: the preview renders the real screen component inside the real frame.

Why this shape

Every version of "have an agent look at all the X" we had tried before, including the pass that found six bugs behind a green pipeline, was a one-off that produced a list and then a second one-off that produced a different list. The five parts exist to make the second run comparable to the first. The surface makes the units the same. The harness makes rendering the same. The probes make the facts the same. The rubric makes the questions the same. The ledger makes the memory the same. Change one of them and you know what moved. Change none and a new finding is news.