3
3 Comments

Had an agent almost drain my API budget over a weekend — here's what actually stopped it happening again

Left an agent running over a weekend to do some scheduled cleanup work. Came back Monday to a bill that made no sense. Turned out a retry loop in my own orchestration code was re-firing the same expensive call every time a downstream service hiccuped — nothing malicious, no prompt injection, just a bug plus an agent that had no concept of "this is too much" baked into it anywhere.

What got me was realizing the system prompt telling it to "be cost-conscious" was doing basically nothing. It's not that the agent ignored it — it's that nothing was actually checking. The limit lived in a sentence, not in anything that could stop a request.

So I rebuilt the whole thing around one idea: the agent should never be the thing deciding whether it's allowed to spend. Something outside it, that it can't see or reason about, has to be the one saying yes or no, every single time, based on an actual running total — not a per-call cap (which is trivially beaten by making more, smaller calls), a real cumulative limit tracked server-side.

Plus a kill switch I can hit instantly without redeploying anything, and a log of every request that I can actually trust wasn't edited after the fact when I'm trying to figure out what happened.

None of this is exotic. It's just applying "don't trust the client" to a client that happens to be an LLM calling your own code — which somehow doesn't feel obvious until you've been burned by it once.

Anyone else building agent stuff hit this? Curious if people are just eating the risk, building this per-project, or if there's a common pattern people have landed on.

on August 16, 2026
  1. 1

    "Don't trust the client" applied to a client that is an LLM is the correct framing, and your point about a per-call cap being trivially beaten is the part most people miss. A per-call limit sets a ceiling on one request and says nothing about how many requests there are, so an agent in a retry loop sails under it all weekend.

    Two things I would add from the same problem space.

    The first is that the cumulative counter only works if it counts the thing that actually costs money. Token-based budgets have a blind spot you already hinted at: an agent can sit comfortably inside its token allowance while making external calls that cost far more than the model calls did. The enforcement has to sit at the action surface, not only at the model layer, or the ceiling is on the wrong quantity.

    The second is that the kill switch and the cumulative limit do different jobs and you need both. The limit protects you from the bug you did not anticipate. The kill switch protects you from the one you cannot diagnose at 2am. A limit alone means the blast radius is bounded but not stopped, and a switch alone means you have to be awake.

    Full disclosure, I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I moved the spend off a meter rather than trying to police it, which is the other honest answer to the problem you are describing. The enforcement pattern still matters for anyone on usage pricing.

    Did you end up making the counter shared across every project, or does each agent get its own budget?

  2. 1

    The difference between a cost limit in a prompt and one enforced outside the agent is pretty striking. Curious how many agent builders have run into this only after seeing an unexpectedly large bill.