1
0 Comments

I built an LLM routing proxy that cuts API costs 60%. You only pay if it saves you money.

I've been building AI-powered features and my OpenAI bill kept climbing. The problem was obvious: I was sending every request through GPT-4o, even simple stuff like text formatting, summarization, or basic Q&A. Probably 60-70% of my requests didn't need a premium model.

I looked for solutions but everything was either "manually pick a model per use case" or required rewriting my application code. I wanted something that just works as a drop-in replacement.

So I built Kestrel. It's a proxy that sits between your app and LLM providers. For each request, an ML classifier analyzes the prompt in under 2ms and routes it to the cheapest model that can handle it. Simple requests go to economy models (Gemini Flash, Groq, Mistral), complex ones stay on premium (GPT-4o, Claude Sonnet).

Integration is one line of code: change your base_url to api.usekestrel.io/v1. Your existing OpenAI SDK code works unchanged.

posted toAvatar for product Kestrel
Kestrel