llmtrace

LLM cost anomaly detection with deploy attribution

Visit Website
May 21, 2026 I built a self-hosted proxy that finds which deploy spiked your LLM bill.

Every team shipping AI features eventually gets a surprise bill. Helicone, Langfuse, Portkey -- they show you that spend went up. None of them tell you which deploy caused it.
  
 I built llmtrace to close that gap. It's a self-hosted Go reverse proxy that:
 - Records every LLM call (cost, tokens, latency, model, prompt fingerprint)
 - Detects per-key spend anomalies on a 7-day rolling baseline
 - Ingests GitHub Actions deploy events  
 - Runs a Gemini agent to name the responsible PR with a confidence score
  
 Live demo: https://llmtrace-681081536857.asia-south1.run.app
 GitHub: https://github.com/Yatsuiii/llmtrace
  
 MIT licensed, single binary, no external dependencies beyond a Gemini API key.
  
 Looking for anyone paying real LLM API bills who'd be willing to try it on their stack. Happy to help with setup.


2 Comments

  1. 1

    This is a very practical wedge. Deploy attribution answers the first incident question: what changed?

    The second question I would keep close to it is: who paid for that change and through which route? For LLM products, the incident trail gets much more useful when the proxy can preserve API key or project, end-user tag, model route, upstream model, retry/fallback path, latency, and the balance bucket that was charged.

    That is the direction we are taking with Tokens Forge: cheaper model access is useful, but the trust layer is the ledger around it. If a deploy switches a workflow from a cheap route to an official/direct model, or a fallback starts firing, the operator should see both the technical cause and the billing semantics in the same place.

    The confidence-score transparency you mentioned is important too. During a bill spike, a boring explanation beats a clever agent answer.

  2. 1

    This is a strong wedge because it ties spend back to a change someone can actually revert.

    One thing I would add early is a very plain "what changed since the last normal day" view: deploy, model mix, average input tokens, average output tokens, request count, and top prompt fingerprints. Founders do not always need perfect attribution first. They need enough context to stop the bleeding and decide whether it was a bug, a model switch, or a real usage spike.

    I would also make the confidence score boringly transparent. If the agent thinks PR 142 caused it, show the 2 or 3 facts that led there. That would make it much easier to trust during an expensive incident.

About

Every team shipping LLM features gets a surprise bill eventually. Existing tools show that spend went up. None tell you which deploy caused it. llmtrace closes that gap.