1
0 Comments

Nobody Asked Who Owned the Agent's Decisions. Until the Agent Got One Wrong.

Why the accountability gap in enterprise AI deployment is not a legal problem — and why fixing it is the deployment team's job, not the lawyer's.


---


The meeting had been going well. Six months into a medical device deployment, the client's VP of Clinical Operations had called us in for a quarterly review. The AI patient triage system was handling 70% of intake routing without human review. Coordinator time was down. Patient assignment accuracy was up. Numbers all pointed in the right direction.


Then the medical director asked a question that changed the energy in the room.


"Two weeks ago the system routed a patient to our low-priority queue. The patient's actual condition required urgent attention. We caught it in time. But I want to understand: when the AI made that routing decision — who authorized it to make that call?"


The room went quiet.


Not because anyone was hostile. Because nobody had a clean answer.


The system had been authorized. Nobody had authorized the specific decisions the system would make.


---


Those are different things. And most enterprise AI deployments — including ours, at that point — treat them as the same.


A deployment gets organizational approval. The procurement process runs. Legal reviews the data processing agreement. IT security signs off on the architecture. The COO signs the contract. The system goes live.


What rarely gets explicit sign-off: the decision scope. What categories of decisions is the AI authorized to make autonomously? What happens when it makes one wrong? Who is responsible for answering the medical director's question?


The enterprise AI conversation in June 2026 has officially moved from "is this real?" to "which part of our organization goes first?" Gartner projects 40% of enterprise applications will have AI agents by end of year. McKinsey reports 62% of organizations experimenting, fewer than 25% at production scale. The gap between those numbers is where this question lives — and it's a question most organizations are not prepared to answer when the agent makes its first consequential mistake.


The agent doesn't have accountability. The model provider has limited liability by contract. The accountability lands on whoever put the system in production. Most deployment teams haven't accepted that explicitly — which means they haven't designed for it.


---


The fertility network: where we learned to write the accountability structure before launch.


After the medical device review, we changed how we scope AI deployments. The change wasn't technical. It was a document.


Before any agent goes live, we now produce what we call a decision authority matrix — a one-page document that answers three questions for every decision category the system will touch:


First, is this decision fully autonomous, human-assisted, or human-required? Autonomous means the agent acts without prompting a human. Human-assisted means the agent produces a recommendation a human reviews. Human-required means the agent cannot take action without explicit approval.


Second, if the agent makes an error in this category, what is the remediation path? Who gets notified, in what timeframe, through what channel?


Third, who in the client organization has signed off that this decision category is appropriate for autonomous handling?


We piloted this format at a nine-location fertility network. The clinical coordinator intake workflow had fourteen distinct decision points. Of those fourteen, we classified six as fully autonomous, five as human-assisted, and three as human-required — specifically, any decision involving a patient flagged with a high-risk indicator, any routing decision that contradicted a physician's standing note, and any action that would contact a patient directly.


The network's VP of Clinical Operations signed off on each classification individually. Not as a block — category by category.


When a patient was misrouted in month three — a real incident, not a hypothetical — the remediation path was already written. The coordinator flagged it within the system, the incident report went to the VP's inbox within fifteen minutes, and we had a documented audit trail showing exactly what inputs the agent had processed and what rule it had applied.


The incident didn't damage the client relationship. The prepared accountability structure was why.


---


What "the agent did it" actually means in a regulated environment.


In the medical device project, the AI triage system operated in a HIPAA-regulated context. We've written before about the vendor assessment that preceded that deployment — the security questions, the data processing agreement, the code provenance reviews.


What we hadn't written about is the conversation that happened after the medical director asked his question.


The client's in-house counsel joined the follow-up meeting. Her concern was specific: under HIPAA's patient safety framework, clinical routing decisions require documentation of the basis for the decision and the identity of the responsible party. The AI had been routing patients. The basis was in the model's inference. The responsible party was — unclear.


We spent three weeks working with their compliance team to produce documentation retroactively: decision logs that traced each routing to the specific inputs and rules the model had applied, a clear statement that all clinical routing decisions remained the responsibility of the VP of Clinical Operations as the accountable human in the system, and a formal amendment to our deployment agreement specifying that we were responsible for the technical accuracy of the system but the client retained clinical accountability for all patient outcomes.


That last point sounds like legal boilerplate. It isn't. It's the answer to the medical director's question. And it should have been written before the system went live, not after an incident surfaced the gap.


In a regulated environment, "the agent made the decision" is not an answer. It's the beginning of a question that needs to have been answered in the contract.


---


WHY ACCOUNTABILITY DESIGN GETS DEFERRED


There's a predictable reason this work gets pushed to after deployment.


The deployment team is focused on getting the system to work. The legal team reviews contracts and data agreements. The compliance team checks regulatory boxes. Nobody owns the question: "when this agent makes a wrong decision, what happens?"


It falls between functions. The deployment team assumes legal handled it. Legal assumes the deployment team scoped it. Compliance assumes both of them addressed it. The agent goes live with the technical infrastructure in place and the accountability infrastructure absent.


We've seen this across every industry we've worked in. Auto retail, medical devices, real estate, film production. The pattern holds regardless of sector. The question doesn't get asked until something goes wrong — and by then, the room goes quiet in the same way ours did.


There is also a sales dynamic that makes this worse. Clients want to hear about what the AI will do. They don't want to hear a forty-five-minute conversation about what happens when it fails. Introducing an accountability scoping session into the deployment process feels like friction, like distrust, like suggesting the system won't work. Every sales instinct runs against it.


We used to skip it. We stopped after the medical director's question, because we realized the discomfort of the accountability conversation before launch is approximately one percent of the discomfort of the accountability conversation after an incident.


---


WHAT THE ACCOUNTABILITY STRUCTURE NEEDS TO COVER


Based on our deployments, a minimum viable accountability structure for any AI agent in production needs explicit answers to four questions — documented, signed, and stored before the system processes its first real decision.


First: what decisions can the agent make without a human in the loop, and what is the explicit basis for classifying those decisions as autonomous? This isn't a technical specification. It's an organizational judgment that a named person in the client organization needs to own. If nobody has signed off on the autonomous decision categories, nobody has actually accepted responsibility for them.


Second: what is the notification and escalation path when the agent makes a decision that turns out to be wrong? "We'll figure it out" is not a path. The path needs to be documented before the first error, because after the first error you're in crisis mode and nobody is designing process.


Third: what audit trail does the system produce for each decision, and who has access to it? In regulated industries, the audit trail is a compliance requirement. Outside regulated industries, it's what determines whether you can reconstruct what happened when someone asks. An agent that can't explain its decisions is an agent you can't defend.


Fourth: what is the human authority structure above the agent? The agent takes actions. Who can override them, in real time? Who can shut the agent down? Who receives escalations? These roles need to be assigned to named individuals, not organizational titles — because when something goes wrong at 2am, the accountability lands on a person, not a job description.


---


ONE THING WE MIGHT BE WRONG ABOUT


The framework above is designed for regulated environments and operational AI — systems making decisions that have direct consequences for real people or real processes. Medical routing, financial approvals, customer communications, supply chain actions.


It applies less directly to analytical AI — systems that produce recommendations, summaries, or insights that a human then acts on. When the human is clearly in the decision loop, the accountability structure is simpler. The agent advises. The human decides. The human is accountable.


The problem is that the line between "advisory" and "operational" blurs faster than most organizations expect. A recommendation that a human technically reviews but almost never overrides is functionally autonomous. The accountability structure of an advisory system should be designed for the behavior that actually emerges in practice, not the behavior that was intended at launch.


We've seen this happen in every deployment where the agent's accuracy was high. High accuracy makes humans trust the system. Trust makes humans review less carefully. Review that is less careful makes the human's role more nominal than real. And a nominal human in the loop is not actually a human in the loop for accountability purposes.


Design the accountability structure for the system you will have in six months, not the system you have today.


---


Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.

posted toAvatar for product Carbuki
Carbuki