If you’re using large language models (LLMs) for classification tasks, you’re likely overspending. TRACER is built to fix that.
👉 https://github.com/adrida/tracer/
Most LLM pipelines send every request to an API—even when the task is simple. TRACER takes a smarter approach.
It learns from your LLM’s outputs and trains a lightweight traditional ML model to handle the majority of cases. The LLM is only used for harder edge cases.
The result: replace 90%+ of LLM calls while keeping performance high.
TRACER doesn’t rely on guesswork. It uses a calibrated “acceptor gate” to decide when the ML model is safe to trust.
You set the desired agreement level (e.g., 95%), and TRACER enforces it. This gives you:
Predictable performance
Controlled risk
Confidence in production
As more data flows through your system, TRACER keeps learning from new LLM outputs. Over time, the ML model improves and handles even more cases.
This creates a powerful loop: More usage → better model → fewer LLM calls → lower costs
TRACER fits easily into existing workflows:
Works with modern embedding pipelines
Supports multiple ML models
Includes tools to analyze performance
It’s ideal for classification-heavy use cases like support tickets, intent detection, and content moderation.
LLMs are powerful—but not every task needs them.
TRACER helps you combine LLMs with traditional ML in a smart way:
Use LLMs where needed
Use ML where it’s enough
Route efficiently between them
TRACER isn’t just an optimization—it’s a better way to scale AI systems.