# Anuj Mulik

> Notes on writing code with AI, at scale.

Anuj Mulik is a software engineer at Featured.com. Insights from working with AI to write code at scale. From Navi Mumbai, India; living in Chandler, Arizona.

The posts here are written with AI, from my own notes, numbers and pull requests. The work, and the mistakes, are mine.

## Writing

- [Mocks first, backend second, glue last](https://anujmulik.com/writing/mocks-first-backend-second-glue-last.md) — The build order we arrived at for shipping product with agents. Perfect the screens on fixtures with no backend. Build every capability against the contract. Wire and test end to end last. Why it did not work before, and what had to change in the codebase to make presentation and logic separable. (October 8, 2026)
- [Every tool you mount is a tax on every turn](https://anujmulik.com/writing/every-tool-you-mount-is-a-tax-on-every-turn.md) — The economics of an agent's authored surface. What an instruction, a tool description and a skill each cost per turn, the budgets that enforce it, the rule that every tool needs a reason that cannot be refuted, and a manifest cut from 23 tools to 7. (October 8, 2026)
- [Read every changelog entry. Then measure it anyway.](https://anujmulik.com/writing/read-every-changelog-entry-then-measure-it-anyway.md) — How we upgrade a framework that releases most days. Enumerate every release between the pin and the target, classify every entry, verify the ones that matter by running them, and treat "undocumented" as a claim that needs evidence. (October 8, 2026)
- [Have an agent use the app and say what hurts](https://anujmulik.com/writing/have-an-agent-use-the-app-and-say-what-hurts.md) — Turning "let an agent test the app and report friction" from a one-off into a system. Five parts, a ledger that remembers, deterministic probes before any model, and the dogfood rounds that feed it. (October 8, 2026)
- [Green CI, six real bugs](https://anujmulik.com/writing/green-ci-six-real-bugs.md) — A 328-file cleanup sweep passed every check. An extra pass asking "what could this break that no test would catch" found six real problems. The questions that found them, and why they are now a standing step. (October 8, 2026)
- [Nobody works past their context limit](https://anujmulik.com/writing/nobody-works-past-their-context-limit.md) — How a long migration was run by a chain of fresh-context worker agents, supervised by a chain of fresh-context overseers, with a cold review before anything touching money shipped. The design, the gates that caught real defects, and the honest costs. (October 8, 2026)
- [Stop asking a language model yes-or-no questions](https://anujmulik.com/writing/stop-asking-a-language-model-yes-or-no-questions.md) — A typed decision model, Jev, answers labels, scores and booleans with a probability for a fraction of a cent. How it unlocked flows we could not afford before, where it does not fit, and the method we use to find the next gain. (October 8, 2026)
- [Does being blunt with the agent work? I asked it.](https://anujmulik.com/writing/does-being-blunt-with-the-agent-work.md) — Five weeks, 218 of my own messages to a coding agent, classified by tone. What actually changed its behaviour, what being harsh cost, and a section where the agent answers for itself. (October 8, 2026)
- [A golden set is not an eval suite](https://anujmulik.com/writing/a-golden-set-is-not-an-eval-suite.md) — How we keep one small labelled set per decision, score it for cents before a merge, and why the set you tuned on is not the number you report. The setup, the traps, and what the live eval suite taught us instead. (October 8, 2026)
- [Delete the comments. The agent believes them.](https://anujmulik.com/writing/delete-the-comments.md) — Code comments were a courtesy to the next human. Coding agents read them as fact, write them as narration, and both go wrong. How we got to a zero-comment budget on new code, and what we do with the context instead. (October 8, 2026)
- [The agent kept a diary. Nobody read it.](https://anujmulik.com/writing/the-agent-kept-a-diary.md) — We tried to make a product observe itself from inside its own code. It cost tokens on every turn and changed nothing. A scheduled agent that runs once a week, outside the product, did what the observer could not. (October 8, 2026)
- [138 database queries to show a home page](https://anujmulik.com/writing/138-queries-to-show-a-home-page.md) — Closing the loop on latency in a chat product. The harness, the eight causes, what each fix measured, and why counts beat milliseconds when the agent is doing the work. (October 8, 2026)
- [I wrote the rules down. The agent broke them anyway.](https://anujmulik.com/writing/guardrails-1-rules-nobody-reads.md) — The first guardrails I put around coding agents were documents. Here is how they failed, what the failures cost, and the shape of the checks that replaced them. (October 8, 2026; Guardrails for coding agents, part 1)
- [When the agent is the maintainer](https://anujmulik.com/writing/guardrails-2-a-repo-whose-maintainer-is-an-agent.md) — Starting a codebase where agents are the primary maintainers from day one. Gate classes, budgets for prompt text, and a rule for when a check is allowed to block. (October 8, 2026; Guardrails for coding agents, part 2)
- [Your test suite is slowing the agent down](https://anujmulik.com/writing/guardrails-3-verify-after-the-push.md) — Moving tests, type checks and evals out of the agent's way, the throughput it bought, and the limits of parallel agents on one machine. (October 8, 2026; Guardrails for coding agents, part 3)
- [The agent made seven decisions without me. I reverted all of them.](https://anujmulik.com/writing/guardrails-4-decisions-need-a-signature.md) — A week of regressions from choices the agent made on its own, and the rules on the human side of the loop that came out of it. (October 8, 2026; Guardrails for coding agents, part 4)
- [I wrote the rule three times. A hook finally held.](https://anujmulik.com/writing/guardrails-5-when-prose-fails-write-a-hook.md) — Some rules protect things that cannot be undone. Those rules do not belong in prose. They belong in a check the agent cannot talk its way past. (October 8, 2026; Guardrails for coding agents, part 5)
- [Stop growing the system prompt. Coach the agent instead.](https://anujmulik.com/writing/a-coach-for-the-agent.md) — We moved behaviour rules out of the system prompt and into one-line, per-turn nudges picked by a cheap judge. Here is what we measured. (October 8, 2026)

## Pages

- [About](https://anujmulik.com/about.md)
- [Contact](https://anujmulik.com/contact.md)
- [Privacy](https://anujmulik.com/privacy.md)
- [Developer docs](https://anujmulik.com/developers.md)
- [Credits](https://anujmulik.com/credits.md)

For agents: [llms.txt](https://anujmulik.com/llms.txt) · [OpenAPI](https://anujmulik.com/openapi.json) · MCP at https://anujmulik.com/api/mcp
