
How a 4-person team discovered $12,000 in untraceable API spend—and what we built to fix it.
Six months ago I did something most founders dread: I audited our API bills.
We were a small team running multiple LLMs in production—Claude for code generation, GPT-4 for customer-facing features. Every developer had their own API key taped into config files. The CTO set a rough mental cap around $300/month. We never looked too closely.
The reality was $3,200/month. $12,000 over the six months we had been ignoring it.
Here's what made it worse: I couldn't tell you which project caused it. Every charge on the invoice was a single lump sum per provider. Was it the recommendation feature? The code review bot? The CI/CD test suite that someone added an LLM step to without telling anyone? I had no idea.
So we did what any engineer-turned-founder would do. We built instrumentation.
The first thing we learned: 90% of our anomaly alerts fired between 2 a.m. and 5 a.m.
Not because we had a bot problem. Because developers deploy late on Fridays, go to bed, and never check whether their new feature has a runaway retry loop. One recursive summarization pipeline—a <code>while True:</code> that re-prompted until the output passed a validation step—burned $800 in a single weekend. We only found it because our Slack channel filled with alerts at 3:07 a.m. and someone happened to be awake.
The second thing: ex-employees are the biggest leak.
Three months after a developer left, his OpenAI key was still active in a staging microservice. Nobody knew it existed. $4,200 in charges over nine months. The key was eventually revoked during a routine audit, but only because our proxy tracks keys by owner and flagged "owner: inactive" automatically. Before that, we never would have noticed.
The third thing: per-project budgets change behavior.
When we rolled out hard caps per project—not warnings, not emails, actual "kill the call at the proxy" limits—something shifted. Team leads started paying attention. They asked for line items. They challenged each other's choices of models. ("Why are you using Claude Opus for that? Gemini Flash is 30x cheaper and the output is indistinguishable for this use case.")
That's a culture change, not just a cost change. And it only happens when attribution is granular enough to make every token spendable feel personal.
After six months of dogfooding, we turned our internal proxy into a product called AiKey. It's open source. It runs locally—no cloud SaaS that can read your keys or prompts. It gives every API call a virtual key tied to a person and a project. Budget caps. Anomaly detection. A single dashboard across Claude, GPT, Gemini, DeepSeek, and any provider with a standard API.
We're a 4-person team. We don't have time to review invoices. We needed something that worked in the background and only interrupted us when there was an actual problem.
If your team is growing its LLM usage and you have no idea who's spending what—this is the tool we wished existed a year ago.
[Try AiKey →] https://aikeylabs.com/zh/i/ih11
Enterprise: aikeyfounder@gmail.com
This matches what we keep seeing when teams move from "one shared provider key" to real production usage. The expensive part is not only the model price, it is losing attribution: which app, which route, which user key, which retry loop, and which fallback path caused the spend.
One pattern that has worked well for us while building Tokens Forge is to treat every LLM call like a payment event: attach owner + project + model route + settlement bucket before the request leaves the gateway, then make caps enforceable at that layer instead of only alerting after the invoice arrives. It also helps to separate "official/direct" routes from lower-cost backup routes, because otherwise teams compare bills without knowing which path actually served the request.
Curious whether your anomaly detection ended up using absolute spend thresholds, token velocity, or model-specific baselines. Token velocity has been more useful for us than daily totals because runaway loops show up earlier.
The payment event framing is exactly right — if you don't capture the attribution context at the moment the request leaves the gateway, you're left reconstructing it from provider logs, which is where the attribution gap starts.
In AiKey we treat each request as a tagged transaction: project, route, feature_id, retry vs. first-attempt, and which provider/key was used. That's the raw data our anomaly detection runs on. The alerts that fired between 2-5 AM were a mix of absolute spend thresholds (e.g., "this route just spent 3× its daily average in 2 hours") and token velocity spikes — the latter caught loops earlier than daily totals would have.
The separation between official/direct routes and backup routes is a nuance we've seen teams struggle with. Without that distinction, you're comparing bills where one side might be 80% backup traffic, which distorts the unit economics. Are you tagging backup routes explicitly in Tokens Forge, or inferring them from fallback chain metadata?