
EchoCache
Save Your LLM Cost By 80% !!
Hey IH — my co-founder and I just launched EchoCache and wanted to share it here. The problem we kept running into: if you're running an AI agent (support bot, internal assistant, anything LLM-backed), most incoming questions are repeats — just worded differently.
"How do I reset my password" and "I forgot my password, help" cost you the same full API call twice, even though you already answered it once. EchoCache sits between your app and your LLM provider (OpenAI, Anthropic, Gemini, or anything OpenAI-compatible). It recognizes when a new question means the same thing as one you've already answered and reuses that response instead of paying for a fresh call.
You just point your app at us instead of directly at the provider — nothing else in your stack changes. For agents handling real support volume, we built a paid plan around this specifically — higher limits, more control over how aggressively it matches — since that's where the cost savings actually add up month over month.
There's also a free tier if you want to try it on smaller traffic first, no card needed. Still early — just the two of us building this. Would love feedback, especially from anyone running a support bot or agent at real volume: what's your repeat-question rate actually looked like, and would something like this be useful for you?
Try it: https://echocache.vercel.app/

Comment