In six B2B AI deployments, compute costs averaged 11% of total project spend. The other 89% doesn't appear on any vendor invoice — and most buyers have no framework for it at all.
The first time a client asked us to break down where their money had gone, we almost didn't want to show them.
Not because the numbers were bad. Because we knew what the conversation would sound like.
They'd budgeted $180,000 for the project. They had a rough mental model: AI API costs, our service fees, some infrastructure. Reasonable assumptions for someone who'd never done this before.
When we walked them through the actual cost breakdown — ours plus their internal team's time — the total came out closer to $340,000. The AI compute costs, the thing most people mean when they say "AI costs," were $31,000.
Nine percent of the real cost. The part everyone talks about.
That project was an outlier only in how clearly we were able to quantify it. In our six deployments, we've tried to track total cost of engagement — not just what clients paid us, but what they spent internally: the IT manager who spent 40% of two months on our integration, the legal team's time on the data processing agreement, the ops manager who was our primary point of contact and whose other responsibilities quietly slipped while they supported our work.
The average compute cost, across six projects: 11% of total. The range: 7% to 16%.
The other 89% distributes roughly like this. Data engineering — cleaning, standardizing, deduplicating, building the pipelines that make the data usable: 35 to 45%. Permission and integration work — the security reviews, the API work, the auth flows, the vendor assessments that nobody scoped: 25 to 35%. Ongoing maintenance and tuning once the system is live: 15 to 20%.
None of this appears on the token bill. Most of it doesn't appear on the service invoice either.
WHY TOKEN COSTS BECAME THE DEFAULT FRAME
This isn't an accident. It's a product of incentives.
AI vendors — the Anthropics, the OpenAIs, the model providers — have a strong reason to make token costs legible. They publish pricing pages. They build dashboards. They send you a bill at the end of the month that shows exactly what you spent, down to the API call. That number is precise, trackable, and designed to be understood.
The rest of the cost structure doesn't have a vendor. Nobody is sending you a monthly invoice for "hours your senior data engineer spent writing normalization scripts for a client's legacy CRM export." Nobody is billing you line-item for the three weeks your IT security team spent reviewing a data processing agreement with a model provider you'd never worked with before. Those costs are absorbed into salaries, into sprint velocity, into the project timelines that slip and the engineers who were supposed to be working on something else.
When one cost is visible and the rest are invisible, the visible one becomes the frame. Not because it's the biggest one. Because it's the only one with a receipt.
There's a secondary factor: vendor demos. Every B2B AI sales process includes a slide or a conversation about ROI, and almost all of them anchor on compute efficiency. Our model costs less per token than the competition. Our architecture reduces API calls by 40%. These are real numbers and they're not dishonest. They're just downstream of the thing that actually determines project economics.
A 40% reduction in API costs on 11% of your total project spend is a 4.4% improvement in total project economics. It's real. It's also almost entirely beside the point.
The real estate staging company: a project that "came in on budget" but cost twice what anyone expected.
A North American property staging company we worked with — asset-heavy, physical furniture, warehouse operations — had a clear budget for their AI virtual staging tool: $95,000, inclusive of our fees and estimated infrastructure.
The project came in on budget. Our invoice was $88,000. API costs were $9,400 — actually below what we'd estimated.
What didn't appear in anyone's budget: their head of product spent approximately 60% of her time over four months on the project. At her fully-loaded compensation, that's $52,000 of internal cost. Their IT team did three weeks of integration work on the inventory management system tie-in — another $14,000. The legal review of the model provider's data terms took their in-house counsel twelve hours — $4,800.
Total actual cost: approximately $159,000. Budget: $95,000. Overage: 67%.
The project was a success. The budget model was fiction.
Nobody lied to build that budget. The client's leadership had simply used the only cost framework they had, which was the one the vendor conversation had made legible: service fees plus API costs. The other $64,000 was organizational overhead that had never been modeled because no one had ever asked them to model it.
The fertility coordination project: where the budget was right but the allocation was wrong.
A different client — a fertility network with nine locations — had done more homework. Their VP of Clinical Operations had run a software implementation before and knew that internal time was a real cost. She'd built a $220,000 total budget, with $40,000 explicitly earmarked for "internal resources."
The $40,000 was the right instinct with the wrong distribution.
She'd assumed the internal cost would be front-loaded: heavy at kickoff, lighter once the system was running. What actually happened was the inverse. The first two months were relatively light — we were doing data work they didn't need to manage closely. Month three was when the integration work started, and that's when IT became a genuine constraint. Month four and five, once the system was in shadow mode, required their clinical coordinator (Maya, the connector we'd met early) to be actively involved in reviewing outputs and flagging edge cases.
The back half of the project consumed 70% of the internal budget. They ran short in month five.
Not a crisis — they absorbed it. But the budget model had been wrong about when internal costs accumulate, which meant resourcing decisions made at month two were made without accurate information about what month four and five would require.
THE THREE COSTS THAT DON'T SHOW UP ON INVOICES
There are three categories of cost that consistently surprise clients and that consistently get underpriced in proposals. They are not new observations — anyone who has shipped B2B AI has encountered all three. What's less common is seeing them modeled explicitly before a project starts.
Data engineering cost. This is the largest single category in our experience, and the most consistently underestimated. It includes: profiling the actual state of the data (what exists, what's missing, what's wrong), normalization work (making data from different systems consistent enough to be useful), deduplication (identifying that the same entity appears in different forms across different systems), pipeline construction (the infrastructure that moves data from where it lives to where the model needs it), and ongoing data quality monitoring once the system is live.
In every project we've done, this work was larger than anyone estimated before seeing the actual data. By a factor of two, on average. The reason is structural: clients know their data is imperfect, but they estimate the imperfection from memory. The actual audit almost always surfaces problems that weren't visible without looking closely.
This category is roughly 35-45% of total project cost. It involves your engineers or ours, and usually both. It produces nothing that looks like AI output until it's done, which makes it politically difficult to resource-justify to stakeholders who want to see the model working.
Permission and integration cost. This is the cost of getting the right people to say yes in the right order on the right timeline. Security reviews, vendor assessments, data processing agreements, API access approvals, IT change management processes, compliance sign-offs. It also includes the actual integration development work: building the connectors between your AI system and the client's existing stack, which in non-AI-native companies almost always involves at least one legacy system with documentation that was last updated in 2016.
The time cost is largely calendar time, not engineering hours — you're waiting for a security team to complete a review, not writing code. But calendar time is a real cost: it delays value delivery, occupies project management attention, and creates uncertainty that makes clients anxious and prompts check-in calls that consume hours from both sides.
This category runs 25-35% of total project cost. A meaningful portion of it is genuinely hard to scope in advance because you don't know the shape of the client's approval processes until you're inside them.
Maintenance and tuning cost. Once a model is live, it requires ongoing attention: threshold adjustments as edge cases emerge, retraining or fine-tuning as the client's underlying data changes, monitoring to catch distribution shift before it affects outputs, and the human review work that shadows the system during the ramp-up period.
This category is 15-20% of total project cost in the first year. It's the cost that most commonly gets omitted from initial budgets entirely, because it's post-launch and therefore feels like a future problem. It isn't. The cost is predictable, and failing to budget for it means the client either absorbs it unplanned or the system degrades without intervention.
WHAT THIS MEANS IF YOU'RE BUYING B2B AI
The budget model you almost certainly built is wrong. Not because you made a mistake, but because the vendor conversation you had was optimized to make API costs legible and leave everything else as a line item called "implementation."
Build your budget from the cost categories, not the invoice structure.
Start with a rough estimate of total engineering hours — yours and the vendor's — from kickoff through six months of live operation. Add a line for internal subject matter expert time (whoever your equivalent of Maya is: the person who knows the business process and the data). Add a line for IT and security review time. Add a line for legal. Add a line for ongoing maintenance at whatever hourly rate applies to whoever will own the system after launch.
Then, separately, add API costs.
If the API cost line is more than 20% of that total, your estimates in the other categories are probably too low.
The buyers who've had the smoothest projects with us are the ones who came into the engagement having already done this math — even roughly. They'd already identified that their senior ops manager would be at 30% allocation for five months. They'd already talked to IT about what a vendor assessment timeline looked like. They weren't surprised by the shape of costs, which meant they weren't surprised mid-project, which meant we weren't managing their expectations in parallel with building the thing.
An accurate budget that's 70% higher than your initial estimate is a better starting point than an optimistic budget that collapses in month three.
WHAT THIS MEANS IF YOU'RE BUILDING B2B AI
If you're pricing your service based on your engineering hours plus a margin, and your clients are absorbing the data and integration costs as unmodeled overhead, you are underpricing and your clients are unprepared. Both of those problems compound over time.
The underpricing shows up in one of two ways. Either you've scoped the data and integration work into your price and clients find it shocking — the proposal is $240,000 when they expected $80,000 — and you lose deals to competitors who quote $80,000 without the honest scope. Or you've scoped to what the market will bear and your team is doing $160,000 of work for $80,000 of revenue, subsidized by engineers who are burning out on work they can't bill for.
Neither is sustainable.
The proposal structure that's worked for us: separate line items for data readiness work, integration development, and model deployment — priced and scoped independently, with explicit client resource commitments attached to each.
Not because clients will always accept it. Some won't, and we lose those deals. But the clients who accept it are the ones who've already thought about what they're buying. They're not treating the data engineering work as a surprising add-on. They're treating it as the first deliverable, which is what it actually is.
The clients who pushed back hardest on that structure were, in our experience, the clients who would have been most surprised mid-project. The proposal conversation is a diagnostic: a client who can't accept an explicit line item for data work is a client who hasn't yet accepted that data work is real.
That's useful information to have before you sign.
ONE THING WE MIGHT BE WRONG ABOUT
The cost ratios we've described — 11% compute, 89% everything else — come from six deployments in traditional industries with legacy data infrastructure. Auto retail, medical devices, industrial manufacturing, real estate. These are industries where data was never designed to be AI-ready, where integration surfaces are old and poorly documented, where security review processes were built for a pre-API world.
For AI-native buyers — modern SaaS companies, tech-forward startups, organizations that already run on clean data infrastructure — the distribution probably looks different. Compute might represent a meaningfully higher percentage of real cost because the data and integration layers are already in place. We don't have clean data on this because we don't work with many AI-native clients.
It's also possible that as tooling matures — better data connectors, standardized integration layers, faster security review processes — the non-compute costs will compress. We'd be happy to be wrong about this. The structure of what we've described isn't inevitable. It's the current shape of the problem in the industries we work in.
What we're fairly confident about: if your cost model for B2B AI deployment is anchored on the token bill, you're optimizing the wrong variable. The month-one invoice from your model provider is not a summary of what you spent. It's a receipt for about one-tenth of it.
Working notes from B2B AI deployment in North America. Part of an ongoing series on what we keep noticing across wildly different industries — and what the industry isn't ready to say out loud.