4
3 Comments

We measured our RAG system in production. It was right 62% of the time. Here's what we did about it.

I'm going to share a number that made our team uncomfortable: 62%.

That was our RAG system's baseline accuracy in production, measured against 150 real user queries, questions pulled from actual support logs, not from a test suite we designed.

38% of the time, users were getting wrong answers. And because the system wrote those answers confidently, in complete sentences, users often didn't know.

We're a small AI engineering team at Ailoitte. We build production AI systems, internal assistants, RAG pipelines, and agent infrastructure for companies that have gotten burned by AI that works in demos and fails in production. This is a post about our own production failure, what we measured, and what we changed.

What we were building

Internal knowledge assistant. One client, ~4,200 documents across Confluence, Drive, and SharePoint. Users ask policy and process questions in natural language.

Standard RAG stack. Good model. Decent chunking. Fast retrieval. Passed user acceptance testing. Went live.

Week three: CTO calls. A VP had used the system for a cross-departmental question, got a wrong answer, and forwarded it to three people. Not great.

The honest post-mortem

We measured everything before changing anything. This is the step most teams skip. They get a complaint, make a change, and hope. Without measurement, you're blind.

150 queries, graded against reference answers. Automated with RAGAS, spot-checked manually.

62% overall accuracy. 41% on multi-document queries (questions requiring information from more than one source). 68% of wrong answers had no hedging language — the model sounded equally confident when it was wrong.

What we changed (in order of impact)

Semantic chunking over a fixed window. Policy documents have logical structures that don't respect 1024-token boundaries. Moving to semantic chunking — splitting on meaning rather than token count — was the highest-impact change.

Hybrid search. We were running vector-only. Added BM25 for keyword matching, fused results with Reciprocal Rank Fusion. Specific document identifiers, regulation codes, and product names that vector search was missing started surfacing correctly.

Cross-encoder re-ranking. Heavy, adds latency, worth it for this use case. Passing the top-20 retrieved chunks through a re-ranker before sending the top-5 to the LLM made a real difference.

Source hierarchy metadata. When two sources contradicted each other, the system guessed. We added authority tags to every document. Primary sources win.

Eval suite, run on every deployment, the discipline change, not the architecture change. You cannot optimize what you don't measure.

Six weeks later

94% overall accuracy. 87% on multi-document queries. False confidence rate dropped from 68% to 12%.

Not because we changed the model. Because we fixed the retrieval layer and built an eval process.

What I'd tell a founder shipping their first RAG system

Build your evaluation suite before you launch. Not after you get complaints — before. Take 100 queries you expect real users to ask. Write reference answers. Automate the grading. These 4-6 hours of setup will save you weeks of debugging production problems blind.

The model is rarely the accuracy bottleneck. If your AI is hallucinating, look at retrieval before you look at the LLM.

And measure in production with real queries, not test cases you designed. The gap between those two accuracy numbers will tell you more about your system than anything else.

Happy to go into more detail on any of the architecture pieces or the eval setup if useful. What RAG implementation challenges are you running into?

→ We're also on Product Hunt if you want to follow along with what we're building: [Product Hunt]

posted toAvatar for product Ailoitte
Ailoitte
  1. 1

    The source hierarchy metadata point is the one people skip most. Everyone reaches for hybrid search and reranking, but 'what happens when two retrieved chunks disagree' is a retrieval-time decision most teams never model explicitly, they just let whichever chunk scores marginally higher win by default. Curious whether you're maintaining that authority tagging manually per source or if it's derived automatically (recency, source type, etc).

  2. 1

    That 62% accuracy hit is honestly terrifying but so relatable! Moving past basic chatbots into real, reliable execution layers always exposes how messy the retrieval layer actually is in production. It's awesome they prioritized semantic chunking and hard evaluations instead of just waiting for a magical new model update to fix it. Building in public with those raw, unfiltered numbers is exactly the kind of transparent reality check that helps everyone ship better products.

    1. 1

      Thanks, Lily! Totally agree, the retrieval layer is where most RAG systems quietly fall apart. Measurement before fixes made all the difference for us. Glad it resonated!