Why the most important work in a B2B AI project happens before anyone touches a model — and why most vendors price it as a free gift.
A manufacturing client came to us with a clear ask: automate their production variance reporting. The current process — a senior analyst pulling numbers from three systems, reconciling them in Excel, and writing a weekly narrative for the ops team — was taking twelve hours a week. They wanted it down to zero.
We said yes. We spent two weeks in their systems before writing a line of code.
What we found: the three systems had three different definitions of "production variance." The CRM counted it from order confirmation. The ERP counted it from shipping. The floor sensor logs counted it from the physical run completing. The twelve-hour analyst process existed, in large part, to reconcile those three definitions every week.
"What does 'production variance' mean to your ops team when they read the report?" we asked the COO.
She paused. "Honestly, it depends on who's asking."
There was no "production variance" to automate. There was a twelve-hour human arbitration process wearing the costume of a reporting task.
Had we started building on week one, we would have automated the inputs from the most accessible system — probably the ERP, which had a clean API. The report would have run in thirty seconds. It would have been wrong in ways that took three months to surface, because the wrongness was quiet: not a crash, not an error, just numbers that looked plausible to anyone who hadn't spent years learning to distrust them.
This is the specific failure mode of skipping the audit. The automation works. The underlying reality it's representing doesn't.
We spent six weeks doing what we now call a process audit before proposing any model architecture. By the end, we had a single-page map of every place a decision got made in the client's ops workflow — which decisions were rule-based, which were judgment calls, which were judgment calls that had been mistaken for rule-based decisions, and which were simply undocumented tribal knowledge that lived in the heads of two people who'd been at the company for a decade.
The client's COO looked at the map and said something we've heard versions of in every deployment since: "We've never seen our own process drawn out like this. It's uncomfortable."
The audit wasn't preparation for the work. It was the first deliverable.
The auto retail group: where the map exposed the real bottleneck.
We've written about the Bay Area auto dealership group before — seven stores, AI voice agent for inbound call handling. The unit economics story closed that deal.
What we haven't written about is the two weeks before we proposed the unit economics.
The sponsor had framed the problem as call volume. "We're missing 30% of test-drive bookings because our staff can't handle peak inbound." Clean problem. Clean AI solution.
Except when we mapped the actual call-handling workflow, we found something the sponsor hadn't mentioned — because he didn't know it: the routing logic for inbound calls was different at each of the seven stores. Three stores had a dedicated receptionist. Two stores were routing to sales staff directly. Two stores had an overflow to a shared service center that the sales managers actively despised and frequently bypassed.
You can't deploy a single AI voice agent across seven locations with seven different routing assumptions. The model would have learned whichever behavior showed up most frequently in the training data and enforced it everywhere. Which meant the AI would have, with high confidence, broken the two stores where the current routing was actually working.
We proposed a three-phase rollout — audit first, standardize the routing logic across stores, then deploy the model. The sponsor pushed back on the timeline. We held the line.
The stores that went live in phase three have a 94% call containment rate. The stores we'd have touched in week one would have had a mutiny.
The medical device company: when the audit finds the project inside the project.
We've also written about the respiratory medical device company — platform architecture, multi-tenant rebuild, HIPAA compliance. The public version of that story is about platform shape.
The private version is about what the audit found.
Their original ask was a UX overhaul. The audit — two weeks of workflow mapping with their clinical team — surfaced something different: their patient-facing app had seventeen distinct onboarding flows, each built by a different contractor at a different time, none of them documented anywhere except in the code. When a hospital system asked "what does your onboarding look like," the sales team would answer based on whichever version they'd most recently demoed. When the implementation team showed up, they got a different version.
The UX wasn't fragmented because of bad design. It was fragmented because nobody had ever mapped the full system in one place.
We didn't start with UI wireframes. We started with a three-week audit that produced something the company had never had: a complete inventory of every patient-facing flow, every backend dependency, and every place a hospital procurement team would encounter a "this is inconsistent" moment during due diligence.
The CEO called it the most valuable document they'd produced in two years of building. It cost them roughly $40,000 in our time.
The AI we eventually deployed — voice triage, automated result interpretation — was built on top of that document. Without it, we'd have been adding automation to a system we didn't understand. With it, we knew exactly which flows were worth automating and which needed to be killed first.
The audit didn't delay the AI. It changed which AI we built.
WHY AUDITS GET SKIPPED
If process auditing is this valuable, why does almost nobody lead with it?
The economics are bad for vendors who don't know what they're doing.
An audit produces a document. A document is hard to sell. You can't demo a document. You can't point to a document and say "this is AI." A client who is excited about AI capability and has already mentally pictured the transformed workflow doesn't want to be told that the first six weeks will be interviews, process mapping, and a report.
Some vendors skip the audit because they believe the AI will surface the issues anyway — the model will train on what exists, the edge cases will show up in production, the client will file tickets, the team will patch. This is true. It's also a description of an extremely expensive way to discover something you could have found in week two.
Other vendors skip the audit because they've priced the project based on what the client expects to pay. If the client expects to pay for "AI deployment," and the audit isn't visibly AI, it gets compressed or eliminated to make the proposal competitive.
We used to do this. We priced audits as a free gift inside the first month. Then we watched three deployments produce exactly the outcomes I've described — not disasters, just slow, complicated, and producing clients who ended up saying "AI is harder than we thought."
The audit was the thing they hadn't paid for. Unsurprisingly, it was the thing that got cut when we were under timeline pressure.
WHAT THE AUDIT ACTUALLY PRODUCES
A process audit in B2B AI isn't a generic consulting deliverable. It's answering four specific questions before any model touches any data.
First: which decisions in this workflow are actually deterministic? Meaning: given the same inputs, would two competent humans always reach the same output? If yes, that decision can be automated. If no — if it's judgment, interpretation, or context-dependent — you're not automating a decision, you're replacing human judgment with machine confidence. That's a different project with a different risk profile.
Second: which parts of the workflow have data, and which parts are invisible? Every business process has formal inputs (the data that lives in systems) and informal inputs (the things that experienced humans know but that never get recorded). Automation can only see the formal inputs. If the decision-making quality depends on the informal inputs — the ones that only Maya knows because she's been there seven years — the automation will perform well on the training data and badly on the cases that actually require judgment.
Third: where does the data mean different things to different teams? The production variance problem is everywhere. "Active customer" means something different to your billing team than to your sales team. "Completed order" means something different to your warehouse than to your finance team. These definitional mismatches are usually harmless in a human workflow because experienced people know to ask. They're catastrophic in an automated workflow because the model picks one definition and enforces it everywhere, silently.
Fourth: what breaks if this process changes? Every workflow has dependencies that aren't visible in the workflow itself — downstream processes that depend on a specific output format, upstream stakeholders who use the current process as a signal for something else, edge cases that the current human handles invisibly. The audit surfaces these before the model breaks them.
These four questions take four to six weeks to answer properly. They cannot be answered by analyzing data alone — they require conversations with the people who actually run the process. That is not a technical limitation. It is the nature of the work.
ONE THING WE MIGHT BE WRONG ABOUT
The audit-first model works clearly when you're deploying AI into a process that already exists — something humans are currently doing that you want to partially or fully automate. The workflow is present. The decisions are happening. The data gaps are findable.
It works less clearly when the client is asking AI to do something genuinely new — a workflow that doesn't currently exist in human form, a capability they don't have at all. In those cases, there's no process to audit. You're building the process and the AI simultaneously.
We've done one project in that category — the AI greenlight scoring system for the film production company. There was no existing "predict which IP to greenlight" process. We couldn't audit something that didn't exist.
What we did instead was audit the decision context: how does the current process work, what signals do experienced producers use implicitly, what does a wrong greenlight decision actually cost. That's a different kind of audit — closer to discovery than process mapping — and it required different questions and a different timeline.
The principle holds: understand before you build. The shape of the understanding changes based on what you're building. For automation projects, that means process mapping. For net-new capability projects, it means decision archaeology. Either way, six weeks of the right conversations is cheaper than six months of building the wrong thing.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.