Why enterprise AI agents go rogue — and why the fix isn't a better model, it's a better spec.
A client called us on a Monday morning. Their AI reporting agent had been running since Friday afternoon.
By the time anyone noticed, it had generated 847 versions of the same quarterly operations report. Slight variations. Different token counts. All saved to the same output folder with sequential filenames. The API bill was $11,400.
We pulled the logs. The agent had completed every step correctly — retrieved the data, ran the analysis, formatted the output. Then it checked its own completion criteria, found something ambiguous, and started again.
There was no bug. There was no hallucination. There was no "done."
This pattern shows up in research that's been making the rounds lately. A production multi-agent system documented publicly ran an infinite conversation loop for 11 days straight, accumulating $47,000 in API charges before anyone noticed. Not because the agents malfunctioned — because two agents kept "working" with no shared definition of completion. The system reported "running" with no error states. From every dashboard, everything looked normal.
The agents weren't broken. They just didn't know they were finished.
MIT research on enterprise AI deployments found that 95% of generative AI pilots fail to deliver promised value. The most cited reason is model quality or data readiness. The reason that almost never gets cited: nobody wrote down what success looked like before the system went live.
Enterprise AI failure is not primarily a model problem. It is, more often, a specification problem — and the most expensive version of that problem is an agent that can't stop.
The manufacturing client where the spec had two parts instead of three.
We worked with a Midwest industrial manufacturer on a procurement workflow automation — matching purchase orders to incoming invoices, flagging discrepancies for human review, routing clean matches to payment. Clearly defined process. Good data.
Three weeks into deployment, the agent started producing what the ops manager called "double approvals" — the same invoice being routed for human review twice, from two different stages in the pipeline.
We looked at the spec. It described the process clearly: retrieve invoice, match to PO, flag discrepancy or route to payment. Two outcomes defined.
What it didn't define: what the agent should do when it had already flagged an invoice for review and that review was still pending. The agent interpreted the pending state as "not yet resolved" — which, technically, was true. So it flagged it again. And again.
The spec had a process description and an output definition. It was missing the third element: a done state. Not "what to produce" but "when to stop producing."
Every AI task spec needs three parts: what to do, what to produce, and what constitutes completion. Most enterprise specs have the first two. Almost none have the third.
The auto retail deployment: where silence was the done state.
When we deployed AI voice agents across a seven-store dealership group, one of our earliest calibration problems was call wrap-up.
After a conversation ended, the agent would sometimes make a follow-up attempt — a check-in call — if it hadn't received an explicit booking confirmation in its post-call window. Reasonable behavior. Except in cases where the customer had asked to "think about it," the agent interpreted the absence of confirmation as an incomplete task and tried again. Some customers received three check-in calls in two hours.
The done state we had defined was "booking confirmed." We had not defined "customer declined to book in this session." Those are two different outcomes, but we'd only written the spec for one.
The fix took forty minutes: add a third outcome to the done state definition — "session complete without booking, no follow-up within 24 hours." The agent had been working exactly as specified. The specification had been working against us.
In B2B AI deployment, the most expensive errors are often not wrong answers. They are correct answers applied past the point where they should have stopped.
WHY DONE STATE IS THE HARDEST PART TO SPECIFY
There is a reason enterprises consistently underspecify completion criteria, and it is not carelessness.
When humans do a task, done state is implicit. An experienced analyst knows when the report is done — not because someone defined it, but because she's done it 200 times and can feel when it's complete. That tacit knowledge doesn't exist in the spec. It lives in the person.
When you automate that task, the tacit knowledge has to become explicit. Every edge case the analyst handled by feel — the ambiguous invoice, the pending review, the customer who said "maybe" — needs a written rule. If you haven't written it, the agent will improvise. And when agents improvise completion, they either stop too early or don't stop at all.
McKinsey research on AI maturity found that fewer than 1% of enterprises have reached what they define as "AI maturity." One of the consistent blockers: the inability to translate operational knowledge into executable specifications. The done-state problem is a specific, tractable instance of that broader failure. It's also why AI governance failures so often start not with the model, but with how the task was defined.
First, treat done state as a first-class deliverable in every AI spec. Not an afterthought. Not something to iterate on in production. Before any agent goes live, write out every state the system can reach, and define explicitly which states are terminal — meaning the agent stops — and which are transitional — meaning the agent continues. If you cannot enumerate the terminal states before deployment, you are not ready to deploy.
Second, distinguish between "task incomplete" and "task unresolvable." These are different states that most specs collapse into one. An invoice that doesn't match any PO is not an incomplete task — it is an unresolvable task that should route to human review and stop. An agent that treats "unresolvable" as "try again" will loop on every exception case in your workflow, which is typically the 10-15% of cases that actually require judgment.
Third, build completion monitoring before you build the agent. One of the consistent lessons from moving AI from pilot to production is that token burn rate is a leading indicator of done-state failure. A properly functioning agent has a predictable cost envelope per task. When cost spikes unexpectedly, the agent has usually either hit an unresolvable state it is treating as "try again," or it has entered a loop between two states with no exit. Monitor cost per task in real time, not monthly. The $47,000 loop ran undetected for 11 days because cost was reviewed on a monthly billing cycle.
ONE THING WE MIGHT BE WRONG ABOUT
The frame above assumes that done state can always be defined in advance — that with enough careful specification, every terminal condition can be written down before deployment.
We are not sure that is entirely true.
In some of our deployments — particularly the AI greenlight scoring system for the film production company — the completion criteria evolved as the system ran. The first version defined "done" as "score generated." After two months in use, the team realized they needed the agent to flag when confidence fell below a threshold, and treat low-confidence outputs as a different terminal state requiring human review rather than a completed score.
The specification was right for what they knew at the time. It needed revision once the system produced outputs they hadn't imagined.
This doesn't undermine the argument for explicit done states. It means done state should be versioned and reviewed, not written once and forgotten. A completion spec is a living document. The $47,000 loop didn't need a perfect spec on day one — it needed someone reviewing whether the spec still matched reality on day seven.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.