Case Study — RAG Development

Fixing a RAG pipeline
that hallucinated confidently.

Illustrative scenario. Composed from patterns we've diagnosed across multiple RAG engagements — not a write-up of one named client project. See our case studies note for why.
Industry: Internal knowledge tooling Timeline: ~5 weeks Related: RAG Development
The problem

The system sounded right, which was the problem.

A team had built a RAG system to answer internal policy and product questions from a growing document set. Early demos looked great — the model answered fluently, in the right tone, citing what looked like the right sections. Then support staff started flagging answers that were confidently wrong: correct-sounding policy language that didn't match what the actual documents said.

The root cause wasn't the language model. It was retrieval. The vector search was returning chunks that were semantically similar to the question but not actually the right source — close enough that the model treated them as ground truth and extrapolated a plausible answer on top. There was no layer checking whether retrieval had actually found the right thing before generation ran.

What it was built on.

A fairly typical RAG setup for this scale of problem — representative of what we see across similar engagements, not necessarily the exact tools in any one project:

Vector database (embeddings-based retrieval) Production LLM API Chunked document ingestion pipeline No retrieval eval layer No groundedness check
The approach

Treat retrieval as something you measure.

Before touching the generation side, we built a golden set — a set of real questions paired with the document sections that should have been retrieved for each. That gave us an actual number for retrieval precision and recall, instead of a vague sense that "it mostly works."

From there:

What changed once retrieval was measured.

This is the pattern of improvement we typically see once a retrieval eval layer goes in — illustrative, not a specific client's measured figures:

Before
Confident wrong answers, no way to measure how often
After
Retrieval precision tracked as a number, re-checked on every change
Ongoing
Low-confidence answers flagged instead of surfaced silently
Related

If this looks familiar.

This is a pattern-level breakdown of what our RAG development work is built to prevent. If your system is giving fluent answers you can't fully trust, it's worth checking whether retrieval quality is actually being measured — not just eyeballed.

Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll tell you honestly whether your retrieval setup has this gap, and what it would take to close it.

Book a free 30-min call → More case studies →