Writing

Mocks first, backend second, glue last

The build order we arrived at for shipping product with agents. Perfect the screens on fixtures with no backend. Build every capability against the contract. Wire and test end to end last. Why it did not work before, and what had to change in the codebase to make presentation and logic separable.

The order we build in now is unusual enough that it is worth stating plainly. First, the screens, every one of them, on fake data, in the real application frame, until the owner picks them from a set of live variants and says yes. Second, the backend capabilities, built against the contracts the screens established. Third, and last, the glue: wiring real data into the screens, and testing the whole thing end to end in a browser.

Agents do most of the typing at every stage. The order is what makes that work.

4
Fixtures per list screen
empty, few, many, long values
73 of 203
Routes in the shell on fixtures
the app's own screens, no backend, no auth
3
Variants per prototype round
each on a named axis, behind a picker, the owner picks
0
Design questions asked in words since
the rule is always a mock

Why this order

Three things are true about building with agents that were less true before.

The agent can produce a complete, working screen on fake data in minutes. It cannot read the owner's mind about which of five layouts is right. So the expensive input is the human's taste, and the cheap input is the agent's throughput. The order puts the human decision where it is cheapest to make, on a screen that exists, before any backend has been shaped around a layout that might be rejected.

The agent is fastest when nothing waits on anything. If the screens establish the contract, the backend capabilities can be built in parallel against that contract, by several agents, none of them blocked on a UI decision or on each other.

And the agent is worst at the last mile: the state nobody specified, the empty tab, the list of two hundred items, the title that wraps. Fixtures make every one of those states a thing you can open, before the data exists to produce it by accident.

Why it did not work before

We tried the order before we had the codebase for it, and it failed in specific ways.

Screens knew where their data came from. A page component read the session, called the database helpers, fetched its own rows, and rendered. There was no screen to put fake data into, because the screen was the fetch. A mock meant a branch in the loader, if (MOCK) return fakeRows, and every such branch is dead code in production and a lie in the mock.

Mocks rotted. A separate mock app, or a component story, drifts from the real screen the week after it is written. The owner's exact words when he asked for a shell of the whole app: all of the UI, none of the backend, dummy data, no auth, local only. A mock app would rot; the shell had to be the app's own screens.

Design questions were asked in words. "Centred column or two columns?" The owner cannot judge a layout from a sentence, and said so: show me options to pick from, I cannot take a call without seeing it first. Every time a question was asked in prose, the answer was wrong half the time and a round trip every time.

States were found one at a time on the live environment. An empty tab here, a dull empty state there, a date in the future on a card about a past event. Each one a separate correction, each one the owner's time. His response: catch everything else so I do not have to point it out one at a time. Things like this are easy to catch on a mock screen where you can simulate all situations and are not dependent on data.

What had to change in the codebase

The decoupling was not a refactor for its own sake. Each change exists because the order needed it.

The page and the screen are two files. The page reads identity and calls loaders. The screen takes the result as props and renders. A screen is a props-only component an agent can draw with no session and no database. The staff console was the first surface split this way, eight pages at a time.

A lint rule enforces the boundary. Components, screens and fixtures may not import the database helpers, the session, the query or command modules, the backend packages, the job runner or the auth library. Client components may still import the server-action stubs, because those are RPC calls in the client bundle, not data access. Reads a page needs live in a queries module named for what they load. Pages that still break the rule are allow-listed until each is split, and the list only shrinks.

Identity is a seam, not a dependency. The layout builds the deferred user promises and renders the session gate itself. The user provider is a pure client component. A pill that shows an unpaid invoice takes its promise from the layout; the sidebar declares the counts it needs and a query satisfies them, rather than the sidebar deriving them.

Fixtures are a registry. A screen name maps to a loader, a set of fixture builders and a chrome. The loader is a dynamic import so enumerating the registry pulls no client bundle into a script. Every fixture builder is seeded from its own name, so a screenshot changes only when the screen does. Every list screen has four fixtures: empty, few, many, and long values. That last one is where titles wrap and numbers overflow.

A preview route renders any screen from any fixture, inside the real application frame when the screen has one. It is open locally and gated by a header everywhere else. There is no component storybook; the preview renders the real screen component inside the real console frame, and the same route is what the automated sweep walks.

The shell runs the whole app on fixtures. One environment variable rewrites every path to its fixture-backed screen, skips the auth gate, and adds a persona switch so the signed-in account is a cookie. A query parameter picks the fixture, another picks a prototype variant. Layouts that read the session got a props-free frame used by both the real layout and the shell. Seventy-three of two hundred and three routes are in it, and every screen added is also a migration step and a sweep unit.

In the chat product the same idea took a different form. New cards are built from a fixed catalog of blocks, never as bespoke components. A scaffold creates the card's data type, fixture and builder and registers the name everywhere the compiler asks for it; it compiles and passes the tests as generated. A satisfies table makes a missing builder or fixture a compile error. A validator and a set of tests enforce composition, spacing on a four-pixel scale, motion rules, and that every piece of copy has an overflow policy, by rendering every demo with long copy and three times the items. A review page shows every card in both themes, in long and crowded modes. A new block needs a design review. A new card never does.

Stage one: perfect the mocks

With that in place, a design round is a set of screens, not a conversation.

A prototype round builds three genuinely different versions of a piece, each on a named axis, layout, density, personality, motion, interaction model, behind a picker, rendered one at a time at full size in realistic context. Three tints of the same idea waste the picker. Every variant fully works: real interactions, real motion, product-shaped copy, no lorem ipsum, no dead buttons. The picker is chrome, not a contestant; its look never adapts to the project. Production code is never touched during exploration. When the owner names a winner, that variant is promoted into the codebase following the project's conventions, and the prototype surface is deleted. The sibling variants go too; git history keeps them.

The owner validates in his own browser. No screenshots are taken for him unless he asks. Once a base wins, the next round diverges around it rather than starting over.

And the agent critiques before showing. Every screen, not a sample, at three widths, in both themes, driven like a first-time user. Data agrees with itself between the list and the card. No developer notes in the copy. The rule came from a round where the owner found six broken mocks in a gallery the agent had glanced at three of.

Two more rules from the rounds. When a production screen's design is already good, the prototype keeps the design and changes only the chrome and the look, and any deviation comes with a stated reason. And when the owner shares an HTML mock, the whole mock is implemented: read its render functions and stylesheet directly, write a parity checklist before coding, port its classes rather than re-deriving them, and screenshot mock against app at the same viewport for every screen and state before saying done. He had to ask four times once. The checklist exists so he does not have to ask again.

Stage two: build the capabilities

Once the screens are signed off, their props are the contract. The backend is built against it, by parallel agents, with a rule that no agent waits on another: build against the spec's contracts and loop back.

The codebase rules that make this safe are about configuration rather than code. Every behaviour table lives in a config.ts beside the code it configures, and a central registry names every lane. A new entity type is added once to one union, and the compiler walks you through every surface that has to handle it, through satisfies-bound tables. Skipping a surface requires an explicit opt-out with a reason. A missing config entry fails loudly, as a compile error where possible and a thrown error otherwise, never a default. Tool failures return a structured envelope the model can see, never an ad-hoc error shape or a client-only toast.

The point of all that for the build order is that the backend cannot quietly disagree with the screens. If a screen expects a field, the type says so. If a capability produces a new kind of thing, every surface is forced to say what it does with it.

Stage three: glue and end to end

Wiring is last, and it is the stage with the strictest gate, because it is where the two halves meet and where the agent's own regressions have lived.

The page calls the real loaders and hands the result to the screen that was signed off on fixtures. Server components load data; the client reads it. Moving a read to the client needs a very good reason stated before building, because the one time an agent did it as a convenience, it added 140 client reads across 53 files and the app got slow for a week.

Then the push gate: type checks, a local production build, and a browser pass over the seven named flows, home through the main views to sign-out. Then the deployment confirmed, and the live timeline script run against the test environment, with reference numbers every change must meet. Then the screen sweep over every registered screen times every fixture, with deterministic probes for overflow, clipped text and small targets, and a rubric for competing actions and placeholder copy. And one invariant became a test in the browser suite: the inbox tab makes zero server calls on load.

What the order bought

It is a migration, not a flip. Most of the app's routes are not split yet, and each new screen pays the split as it is touched. But the direction is set by a lint rule and a registry rather than by intent, and that is the difference between an order we follow and an order we talk about.