Blog

How we ship software with AI agents without losing control

3 min read

  • agents
  • engineering
  • process

AI agents now write much of the code we ship. That sounds like a loss of control. In practice it has made our work more controlled, because an agent only does well inside the structure a careful team would want anyway.

Here is the loop we use on our own products and on client work.

Start with a written spec

Nothing starts from a chat message. Every change begins as a short specification: the outcome, who it is for, the constraints, and the acceptance criteria that will tell us it is done.

The spec is for people first. Writing it forces the decisions that are easy to skip: what happens on an empty screen, what the error message says, what we are not building. Once it is written down, an agent can read it as often as it needs, and so can the reviewer later.

Turn the spec into small tasks

The spec becomes a plan. Each task is small enough to finish and test on its own. It names the exact files it touches, the interfaces it depends on and the tests that must pass.

This is the step that makes agents reliable. A vague task invites a model to fill gaps with guesses. A task that says “add this function to this file, with these three tests” leaves little room to wander, and when something goes wrong the damage stays inside one task.

One agent per task, with a narrow scope

Each task goes to an agent that sees only what it needs: the spec, its task and the relevant code. It writes the tests first, then the implementation, and runs the tests itself.

Scope also means permissions. Agents work in the project they were given. They do not run destructive commands, push to shared branches or touch production without a person approving it. Every task leaves a trail: what was asked, what changed and what the tests said.

Review by someone other than the author

An agent is a poor judge of its own work, in the same way a tired engineer is. So nothing is accepted on the author’s word. A separate reviewer, starting from fresh context, reads the change against the spec and the plan and reports what is missing or wrong.

The reviewer gets specific questions. Does the change do what the task says? Does it do anything the task did not ask for? Do the tests check behaviour, or do they just restate the code? A failed review sends the task back and the loop runs again.

Verify in the real thing

Passing tests are necessary but not enough. Before we call work done, we check it where people will meet it. For a website that means a real browser at phone and desktop widths, with accessibility and privacy checks. For a backend it means calls against a running service, not only mocks.

This step catches bugs that unit tests never see: a button that exists but sits under the header, a layout that breaks at one width, a page that shows a detail it should not.

Keep the knowledge outside the model

Agents forget between sessions, so we do not ask them to remember. Our projects, their dependencies and who owns what live in a shared hub that every session reads first. Recurring work is written down as agent skills, more than 40 of them now. Each one is a reviewable playbook, not a clever prompt someone typed once.

When a skill turns out to be wrong, we fix the skill. Every later run gets the fix.

Where people stay in charge

People decide what to build, approve the spec, approve anything that cannot be undone, and settle it when the reviewer and the author disagree. Agents propose, implement and check.

This is ordinary engineering discipline. What changed is that it now happens every single time, because the process is written down and the agents follow it without getting bored or cutting corners on a Friday. That consistency is why we trust the output enough to ship it.

Work with usNeed something like this built? Say hello