Have an agent use the app and say what hurts
Turning "let an agent test the app and report friction" from a one-off into a system. Five parts, a ledger that remembers, deterministic probes before any model, and the dogfood rounds that feed it.
The request sounded simple: have an agent use the app the way a person would and report the friction. We had done it once, by hand, with a script that rendered every email template and an agent that looked at them. It found things. It also had no rubric, no memory of what it found last time, and no way to tell run two from run one. The second time the same request came back with a different noun, screens instead of emails, we made it a system.
The shape
A sweep is five parts. An instance names all five; shared plumbing makes an instance a few files.
surface → harness → probes → rubric → ledger
- Surface: an enumerable list of units, from code, never from clicking around. Screens times fixtures times viewports. Templates. Personas times tasks. Enumerable is what makes run N comparable to run N+1.
- Harness: renders one unit deterministically with no session and no database. For screens it is a preview route over a fixture registry. For emails it is the template renderer with its preview props. For flows it is a seeded account on a preview deployment. The harness is the expensive investment and the one every sweep shares.
- Probes: deterministic extractors over a rendered unit. Horizontal overflow, text clipped with no tooltip, controls under 24 pixels, accessibility violations, page errors, placeholder text like
undefinedorNaNor lorem ipsum in visible copy, dead links, generic link text, an email over the size at which a mail client clips it. No model. They never lie, so they can gate. - Rubric: a fixed, versioned question set for a typed judge over the probe facts. Booleans and small enums only, referencing only fields in the state. "Is the design good?" is not a question. "Are there competing primary actions?" is. An answer becomes a finding only above a fixed confidence, so a score drifting across the midpoint does not open and close the same finding every other run.
- Ledger: a tracked file of findings per sweep. Each finding has a stable id, a first-seen and last-seen run, and a status of open, accepted or fixed, with a reason for accepted ones. A run reports new, regressed, still open, fixed and accepted. Never a fresh list of eighty.
The rules that make it a system
A sweep with no ledger is a demo. If run two cannot say what changed since run one, it is not done. The ledger is tracked in git, so a pull request that fixes a finding shows the finding leaving the file, and a reviewer can see it.
Probes first. If a deterministic check can answer it, do not ask a model. The judge is for what the DOM cannot say. A person or a vision model is for what the judge cannot see.
Only units the run swept can close a finding. A partial run, restricted to one screen, is therefore safe. It cannot accidentally mark the rest fixed.
Accepted findings stay in the ledger with a reason, so a nit is not re-litigated every run. They close like any other when they disappear.
Fixtures over mocks. The harness renders from props. There is no if (MOCK) in a data loader anywhere. That rule turned out to be the one with the largest consequences, and it gets its own post.
Changing a rubric question is a pull request. So when a score moves, the product moved.
Three sweeps
Emails. Every template rendered with its preview props, checked as markup, opened in a browser at 720 and 375 pixels for overflow and small targets, run through the accessibility checker, then the email rubric: how many calls to action, whether their labels are generic, whether the subject matches the body, whether the opening says why the reader got it. The first run found the real root causes behind a dozen findings: a link colour with insufficient contrast in the shared template, and a box-sizing rule in the body. Fixed once, closed everywhere.
Screens. Every registered screen times every fixture times two viewports, drawn by the preview harness inside the real application frame. Probes measure the screen's own content, not the frame around it, so a navigation defect is not reported once per screen; the frame is registered as its own unit and swept once. The screen rubric asks about competing primary actions, unexplained internal shorthand, placeholder copy, and whether an empty state says why it is empty and what to do next. Adding a screen to the registry adds it to the sweep.
Flows. The one an agent drives by hand. Personas and tasks are the surface, with tasks worded as the persona would think them, never as routes. The agent signs in as the persona on a preview deployment, attempts each task using only what the persona would know, and journals every moment of friction with a type, hesitation, dead end, surprise, unclear result, unreachable, and a key naming the thing rather than the moment. The journal is then triaged through a friction rubric that answers two questions, who would hit it and what it costs, and reconciled into the same ledger. Tester-only entries are dropped. Meant to run weekly as a scheduled agent.
What the dogfood rounds taught
Before the sweeps there were dogfood rounds: five functional agents in parallel, each on one surface, plus a latency agent, plus a human pass with real prompts. The process document that came out of round one is mostly a list of what went wrong, which is the useful kind.
- Plan for a verification round. Round one's agents reported four to six false positives or mischaracterised issues. Resolving them took thirty minutes and had to happen before any implementation tickets were cut, or the tickets carried the false positives forward.
- Seed the data first. Three of five agents were partially blocked by an empty workspace: no profiles to open, no active instances to rename, no rows to paginate. A seed script that creates one of everything, and a reset that tears it down, now runs before any round.
- Guard rules in every brief. Do not claim a feature is broken without retrying the click once after a second, because portalled menus mount slowly. Check the feature flags before reporting a missing feature. Distinguish a headless-browser artefact from a real bug; anything involving speech, media devices or hardware needs a real-browser retest. When reporting performance, separate cold from warm and sample the mutation stream, not just the first mutation.
- One long-lived tracking issue, not a fresh issue per round. Each round posts a structured comment: scope, trigger, new bugs, regressions, verified fixed. Deltas over time are what you want to read.
- Separate dev-server cost from real cost. Round one measured 2.6 to 3.1 second cold loads that were almost certainly the dev bundler compiling. Once a quarter, run the cold sweep against a production build.
What it does not do
No golden-screenshot pixel diffing: the ledger is about findings, not pixels, and screenshots are artefacts for a reviewer. No vision model in the loop: when one is wanted, it is a generator in front of the rubric, writing the same journal the flows sweep uses, and triaged the same way. No component storybook: the preview renders the real screen component inside the real frame.
Why this shape
Every version of "have an agent look at all the X" we had tried before, including the pass that found six bugs behind a green pipeline, was a one-off that produced a list and then a second one-off that produced a different list. The five parts exist to make the second run comparable to the first. The surface makes the units the same. The harness makes rendering the same. The probes make the facts the same. The rubric makes the questions the same. The ledger makes the memory the same. Change one of them and you know what moved. Change none and a new finding is news.