Almost every time I saw things go wrong with a coding agent, the problem was not the model. It was me jumping straight to code without writing the spec first. The agent fills every gap with a guess, the guesses pile up, and by the third file you are reviewing something you never asked for. This site, the desktop apps and the content engines I publish were all built the same way, so I am going to show you the loop exactly as it runs here today.
Why spec first
The spec is the cheapest place to be wrong. Before writing any code, I write one short document saying what will exist when the work is done: the routes, the components, the copy keys, the tests that must pass, and what stays out. It usually fits on one screen. It is not bureaucracy, it is economy: the agent reads that document at the start of every session and begins at the level of someone who knows the project, instead of trying to work it all out again from the file tree.
The spec also settles arguments before they happen. If the design is black and white, the spec says so, and no agent comes along proposing an accent color at two in the morning.
Splitting the plan into tested tasks
The spec becomes a plan, and the plan becomes numbered tasks. I keep each task small enough to finish in one session, and each one carries three things: the files it may touch, the acceptance criteria, and the test that proves them. If a task has no test, it means I will have to check it by hand later, so it does not even enter the queue.
The check is always the same command, and it is the only definition of done:
npm run check && npm run build
Lint, types, tests, then the production build. If that line is green, the task is done. If it is red, it is not done, however good the diff looks.
Agents in parallel, each in its own worktree
A task that shares no files with another can run at the same time. I give each agent its own git worktree and its own branch, born from the current main, so nobody edits the same file and nobody waits. One takes the hero, another takes the copy, a third takes the footer. They all get the same spec, the same commands and the same rule: only touch the files the task names.
When a task finishes, the branch rebases onto main and goes in. Merges are one at a time and small. That discipline alone ended almost every conflict I used to have.
Review gates
Every merge goes through the same gates. The command above has to be green with the output pasted, not promised. A screenshot proves what the browser shows, because a passing test does not mean the page is right. And the reviewer, whether it is me or an agent, reads the diff against the task, not against the whole repository. If the diff touched a file the task did not name, it goes back.
What breaks
Three things that broke on me in this site, and I would rather tell you than let you find out alone.
Hydration and reduced motion. A hook that reads prefers-reduced-motion on the first render produces one HTML on the server and another on the client, and React complains. I fixed it by rendering the static version first and only switching after mount.
Symlinked node_modules with Turbopack. Sharing dependencies between worktrees through a symlink looked clever and broke the dev server in a way the error messages never explained. Now each worktree installs its own.
Screenshot with no page open. An agent asked the browser for a screenshot before opening any page, got a blank image and reported the layout as fine. The rule became: open, wait, take the screenshot, then look at the file with your own eyes.
Closing
None of this is about trusting the model less. It is about giving it what any good teammate would want too: a clear spec, a task you can actually finish, a test that says when it is finished, and a review that reads what it really did. Do that, and you can let it run.