Most teams running AI API costs hit the same wall around month 4 — the bill quietly 2-3x's as soon as the product gets real traffic. We hit it last quarter and went deep on what actually compresses it.
The short version: caching in AI gateways is not one feature — it's two. And most aggregation platforms (OpenRouter, Portkey, even newer ones) only ship half by default. That half saves real money. The other half, stacked on top, is what cuts the total bill roughly in half.
Full technical breakdown: [AI Gateway Caching 2026 → tokenmix.ai](https://tokenmix.ai/blog/ai-gateway-caching-l1-l2-guide-2026)
The two layers in indie-hacker terms:
- L1 — result cache: skip the model entirely on duplicate or semantically similar requests. 100% savings per hit. Helicone does this via a 1-line base_url swap. Self-hosted Redis works too.
- L2 — prompt cache: the vendor (Claude, OpenAI, DeepSeek, Gemini) caches the KV state of your stable prompt prefix. Model still runs, but input cost drops 50-90%. Claude reads cache at 10% of input price (90% off). DeepSeek same. OpenAI auto-caches anything ≥1024 tokens.
Most teams get only L2 because vendors auto-enable it. L1 is what gateways typically skip.
Real math for 10M requests/month (Claude Sonnet 4.6, 4K avg input tokens):
- No caching: $195K/month baseline
- L2 only (80% prefix hit rate): $119K/month (−39%)
- L1 + L2 stacked (25% L1 hit + L2 on rest): $90K/month (−54%)
Savings are compound, not additive — the requests L1 absorbs never touch L2, so L1 savings don't dilute.
4 things we learned running this in production at TokenMix.ai:
1. Prefix stability is everything for L2. Middleware that rewrites system prompts (adds timestamps, user IDs, etc.) kills the cache key. Pinning our system prompt format alone 4x'd our real cache hit rate.
2. Semantic cache false positives hurt. We set cosine threshold at 0.88 early and got some wrong answers. Moved to 0.95+, accepted lower hit rate for correctness.
3. Vendor lock-in is worse without gateway caching. Direct Claude integration gets L2. But migrating to another model means rebuilding cache infra. Gateway-level caching makes the switch cheap.
4. Measurement is not optional. No cache-hit-rate dashboard = flying blind. Shipping this as a first-class metric in TokenMix.ai's usage panel.
If you're early: just make prompt prefixes stable so L2 works when you scale. Don't build L1 yet.
If you're already at scale: the full write-up with pricing tables, code examples, and architecture patterns is on [our blog](https://tokenmix.ai/blog/ai-gateway-caching-l1-l2-guide-2026) (~2500 words), tighter dev version on [Dev.to](https://dev.to/tokenmixai/ai-gateway-caching-explained-why-l1-l2-cache-layers-cut-90-of-your-llm-bill-45ab).
Happy to answer questions on L1 semantic cache tuning in comments — that's where most teams screw up first.
cache really save my money