Kestrel

Route smarter, spend less

Visit Website
April 7, 2026 I built an LLM routing proxy that cuts API costs 60%. You only pay if it saves you money.

I've been building AI-powered features and my OpenAI bill kept climbing. The problem was obvious: I was sending every request through GPT-4o, even simple stuff like text formatting, summarization, or basic Q&A. Probably 60-70% of my requests didn't need a premium model.

I looked for solutions but everything was either "manually pick a model per use case" or required rewriting my application code. I wanted something that just works as a drop-in replacement.

So I built Kestrel. It's a proxy that sits between your app and LLM providers. For each request, an ML classifier analyzes the prompt in under 2ms and routes it to the cheapest model that can handle it. Simple requests go to economy models (Gemini Flash, Groq, Mistral), complex ones stay on premium (GPT-4o, Claude Sonnet).

Integration is one line of code: change your base_url to api.usekestrel.io/v1. Your existing OpenAI SDK code works unchanged.

Comment

About

I was building AI features and realized I was sending every request through GPT-4o, even simple ones like reformatting text or answering basic questions. My API bill kept climbing but most of that spend was wasted.