
Kestrel
Route smarter, spend less
I've been building AI-powered features and my OpenAI bill kept climbing. The problem was obvious: I was sending every request through GPT-4o, even simple stuff like text formatting, summarization, or basic Q&A. Probably 60-70% of my requests didn't need a premium model.
I looked for solutions but everything was either "manually pick a model per use case" or required rewriting my application code. I wanted something that just works as a drop-in replacement.
So I built Kestrel. It's a proxy that sits between your app and LLM providers. For each request, an ML classifier analyzes the prompt in under 2ms and routes it to the cheapest model that can handle it. Simple requests go to economy models (Gemini Flash, Groq, Mistral), complex ones stay on premium (GPT-4o, Claude Sonnet).
Integration is one line of code: change your base_url to api.usekestrel.io/v1. Your existing OpenAI SDK code works unchanged.
About
I was building AI features and realized I was sending every request through GPT-4o, even simple ones like reformatting text or answering basic questions. My API bill kept climbing but most of that spend was wasted.

Comment