Running a 24/7 AI software factory: builders and one advisor

Published Updated 12 min read

One AI coding session can finish a ticket. It cannot clear a backlog of several hundred. Heavy workloads like that come with a deadline and a bar for proof that cannot be lowered to save time. One engineer cannot do that alone, and neither can one agent.

So we stopped treating the agent as an assistant and built a small company out of agents. A handful of builders, one advisor, one human owner. It runs around the clock. This post is how it is set up, how a ticket moves through it, and what breaks.

Key takeaways

  • Give each builder agent its own area of the codebase, its own goal file and its own ledger.
  • Put one advisor agent in charge of routing, merging, deploying and closing tickets. It writes no product code.
  • Keep all state in files. Sessions die, and a new one must be able to boot from disk alone.
  • Ship in release trains: many tickets, one pull request. Drop a failing ticket instead of debugging inside the train.
  • The real ceiling is model usage limits and disk space, not CPU.

The problem

Heavy work usually arrives with three things that cannot move:

  1. The deadline. A large backlog has to be live by a fixed date.
  2. The proof. Each ticket has to be built test-first and carry its own evidence that a human reviewer would accept.
  3. The safety rules. No real user data on the development server, no shortcuts around checks, and some decisions that only a named person could make.

A single long agent session fails here in a predictable way. It loses track of what it promised, it waits when it is blocked, and when it crashes the plan goes with it. More sessions without structure are worse: two agents edit the same file, and nobody knows which tickets were dropped.

What “factory” means here

The term is recent. In early 2026 the StrongDM team described a “Software Factory” where agents write, test and ship code, and people write the specs and watch the scores. Others call the far end of this a “dark factory”, after plants where robots work with the lights off.

Ours is not dark. People set the goals, own the risky decisions and can read every ruling. We use “factory” for a simpler idea: a standing team of agents that takes tickets in at one end and puts tested, deployed changes out at the other, all day.

The shape of the factory

Picture a small software team where everyone is an agent except the owner.

Role Does Never does
Builder (several) Builds tickets test-first in its own area of the code, commits locally, reports each finished ticket Pushes, merges, or edits another builder’s area
Integrator (two builders take turns) Bundles finished tickets into one release train and opens one pull request Debugs a failing ticket inside the train
Advisor (one) Rules on engineering questions, routes work, merges, deploys, verifies, posts evidence, tracks every ticket Writes product code, or asks the owner about engineering choices
Owner (a person) Sets the deadline and the limits, decides values only a person may decide Reviews every change by hand

Each agent is a full, independent Claude Code CLI session on the same machine. That matters. A builder has the whole tool: its own context, its own permissions, and the power to spawn its own subagents and parallel worktrees when a ticket is bigger than it looked. They talk through the built-in session messaging, and every message has a fixed shape. A builder that finishes a ticket sends the failing-test commit, the passing commit, the test count and the exit code. A builder that is stuck sends a numbered question with options and its own recommendation, then moves to other work instead of waiting.

The owner hears from the advisor about three things only: a business value nobody has decided, an act that legally needs a human, and anything that touches real user data or production. Everything else is the advisor’s job to decide.

How one ticket moves

Ticket
  -> advisor reads it fully, rules on open questions, adds a row to a builder's ledger
  -> builder: failing test -> passing test -> break the fix on purpose (the test must catch it)
              -> targeted tests on a remote runner -> local commit
  -> builder reports the ticket as finished, with commit ids and a receipt
  -> integrator cherry-picks finished tickets into a release train: one push, one pull request
  -> required checks run
       a red ticket is dropped and re-routed; a flaky check is read, then re-run
  -> advisor merges when every check is green, watches the deploy,
     and reads the server's health endpoint until it reports the merged commit
  -> advisor posts evidence on each ticket and moves its state

Two details matter more than they look.

A ticket’s evidence is its own. “The train passed” is never proof for a ticket. Each one carries its own failing and passing commits and the tests that cover it.

The health endpoint reports the running commit. A green deploy job is not the same as the new code being live. The advisor reads the version back from the server before it closes anything.

The files that hold the system

Agents forget. So the factory lives in a shared folder of plain files, and a session is only the thing that reads and writes them.

  • A goal file per builder. Its scope, its priority order and what “done” means.
  • A loop file per builder. What it checks on every cycle.
  • A ledger per agent. One dated row per piece of work, with the ticket id, the batch and a status word. Rows are only ever added.
  • One file per ruling. Each has an id, so a builder can cite it instead of asking again.
  • An evidence draft per ticket, written by the builder and posted by the advisor.
  • A handoff file. A “live state” block that a brand new advisor session reads first.
  • The advisor’s operating manual, written as an agent skill and re-read on every run.

How we set it up

  1. Write the advisor’s manual first. Its prime goal, the checklist it runs every cycle, the rules for trains, how it decides, what it is allowed to touch and a list of things it must never do. Keep it dense, because it is read every run.
  2. Split the codebase by seam. Give each builder an area where it will not edit the same files as anyone else. Payments in one, search in another, and so on. This step decides how many builders you can have.
  3. Create the ledgers before any work starts, and make “no row, no work” a rule.
  4. Install one build discipline for every builder: failing test, passing test, a deliberate break to prove the test bites, targeted tests, evidence.
  5. Move heavy work off the workstation. Type checks and test runs go to a remote runner. The machine running six sessions has to stay responsive.
  6. Protect the main branch with required checks, and end the deploy with a health endpoint that returns the running commit.
  7. Give the advisor an app identity in the ticket system, not a personal key. Only the advisor writes to tickets. Builders get what they need copied into files.
  8. Start the builders, each pointed at its goal file.
  9. Start the advisor on a timer. Ours runs several times an hour. Add monitors that fire when a train’s checks change or the server’s version changes.
  10. Schedule the full test suite at night, on clean clones. It never runs inside a push or a train.

The advisor’s loop

Every cycle, in order: check where the current train is, compare the pace against the deadline, check disk and load, then run a roll call. The roll call is the same five questions to every builder: what are you on, what is finished but not yet routed, what are you stuck on, what are you waiting on me for, and how many parallel lanes are you using.

Then it answers every blocked item in that same cycle. If a builder is waiting on the advisor, that is logged as the advisor’s failure.

One rule kept tickets from vanishing. Any ticket the advisor touches and leaves unfinished must have an open row in someone’s ledger. A small script checks every recently touched ticket, and it has to print zero orphans.

What went wrong

This is the part worth reading.

  • Sessions die. A usage limit can freeze every agent at once, and killed sessions come back under new names. Because the state was on disk, they re-read their ledgers and carried on. The advisor now looks up session names on every run and never hard-codes them.
  • Test databases drift. A push gets refused when a builder’s private test database still held a schema change from a ticket we had pulled out of the train. The fix was to create a fresh database after every drop. We never reset one.
  • Runners flake. A required check went red because the hosted runner could not create a temporary file. Every test but one had passed. Reading the actual log saved a good ticket from being dropped.
  • Dropping a ticket can remove a safety check. Reverting one ticket also removed guards that ticket had added, and a gate refused the push. The removal had to be declared, with a reason and a plan to bring it back.
  • The lock was not a queue. Commits waited far too long for a shared pre-commit slot, because it was a polling race that later arrivals kept winning. It became a first-in, first-out queue.
  • The advisor was wrong. It told a builder to keep two tickets in a train. The builder’s measurement showed otherwise, and the ruling was withdrawn within the hour. It also stated with confidence that a server setting was causing a bug. A read-only check showed the setting did not exist. Both became written rules: measure before ruling, and change position only on new measured facts.
  • A helper printed a secret into a local log. Nothing left the machine. It still produced hard rules: capture a token into a variable and never print it, read environment files by key name only, and rotate anything that was exposed.

The rules that kept it safe

  • No builder pushes, and no builder skips a hook.
  • A red ticket is dropped to the next train. A train is never held to debug one ticket.
  • Test databases are created, never reset or dropped.
  • No real user data on the development server. Automated walk-throughs use synthetic accounts and are labelled as not a human decision.
  • A business value nobody has decided ships as a setting that fails closed. The software works and refuses safely, and only the value waits for a person.
  • Lists of allowed exceptions only shrink.
  • Agents never kill processes by name, because every session runs as the same user.

What to expect

A factory like this runs continuously and closes dozens of tickets a day. A train takes minutes to push and a little longer to pass its checks. The advisor writes many rulings to file in a day and reverses some of them, which is the system working.

More rules do not make it faster. The levers are more parallel trains, every lane full, and more model capacity.

What we would do differently

  1. Plan for usage limits on day one. They are the real throughput ceiling. We put workers on separate Claude Code accounts, one seat per worker, so a limit pauses one builder and the rest keep going.
  2. Make test databases disposable from the start. Every polluted database costs real time.
  3. Make every check say why it is red in plain words. We lost time decoding what a failing count actually measured.
  4. Build the orphan check into the act of touching a ticket, not a later audit.
  5. Wrap any script that can print a secret, so no agent can run it bare.

Why plain Claude Code, and not a framework

We looked at the tools built for this before writing our own rules.

Option What it is Good at Why we did not use it here
Paperclip Open-source server and dashboard that runs agents as a company: org chart, heartbeats, budgets Cost control per agent, one dashboard, works with any agent runtime A second system to host and learn. Its unit is a company. Ours is one repository and one backlog
Hermes Agent Open-source personal agent from Nous Research that runs on your server and writes its own skills Long-lived memory, learning from experience, chat apps as the interface Built as one personal assistant, not a team of builders on one codebase. Broad permissions need a sandbox
Gas Town Steve Yegge’s orchestrator: a “Mayor” agent directs many worker agents, with tasks stored in git Scale, 20 to 30 agents at once A large codebase and its own vocabulary to learn before the first ticket ships
Claude Code agent teams Built in: a lead session spawns teammates that message each other Nothing to install, good for one task split across a few agents Still experimental, and a team lives for one task. We needed builders that run for days
Claude Code Projects New in beta: a coordinator splits a goal across parallel cloud sessions with shared memory Runs in the cloud with the laptop closed, each thread on its own branch Beta and still rolling out when we built this. Our checks needed local databases and tools

What we run is the plain version: ordinary Claude Code sessions, the built-in messaging between sessions, agent skills for the manuals and files for the state.

What that buys you

  • Nothing extra to install, host or pay for.
  • Every worker is a complete Claude Code CLI, so it has full power and you have full control over what each one may touch.
  • Workers fan out on their own. One builder can run several subagents at once.
  • Each worker can sit on its own Claude Code account, so usage is spread out and one limit does not stop the whole factory.
  • Every rule is a text file your team can read and change.
  • It works on your machine, with your databases, your checks and your access.
  • As the native features mature, the same files and rules move across. Projects and agent teams replace the plumbing, and the discipline stays.

What it costs you

  • No dashboard. You read ledgers, not charts.
  • No built-in budget control. Usage limits are the only brake.
  • It is tied to one model vendor.
  • You write and maintain the rules yourself, and the first version will be wrong in places.

If you want budgets and an org chart across many kinds of agents, look at Paperclip. If you want one agent that learns your habits, look at Hermes. If you want a team of coding agents on one repository this week, plain Claude Code is enough.

Where this fits

This is the same discipline we described in how we ship software with AI agents, scaled from one agent per task to a standing team. The spec, the test-first build and the independent check do not change. What changes is that coordination becomes a job of its own, and that job needs its own agent, its own files and its own rules.

If you have a backlog and a deadline and want this set up around your repository, that is our AI software factory service.

FAQ

What is an AI software factory?
A group of AI coding agents that work in parallel on one codebase, each with its own area, plus a coordinating agent that routes work, merges, deploys and keeps every ticket accounted for. People set the goals and the limits.
How many agents can work on one repository at once?
We run a handful of builders, each with a few parallel worktrees. The limit is not the machine. It is model usage limits, disk space and how cleanly the codebase splits into areas that do not share files.
How do you stop AI agents from breaking production?
Builders never push. Every ticket needs a failing test, then a passing one, then its own evidence. Changes ride in batches through required checks, and only the coordinating agent merges and deploys. Anything touching real data or production needs a person.
What happens when an agent session crashes?
Nothing is lost if the state lives in files. Each agent keeps a dated ledger of its work, so a fresh session reads the ledger and a handoff file and carries on.
Do you need a framework like Paperclip or Gas Town to run an AI software factory?
No. Ours uses plain Claude Code sessions, the built-in session messaging, agent skills and files in a shared folder. Frameworks add dashboards, budgets and scale. They also add a second system to run and learn.

Work with usNeed something like this built? Say hello

Related posts