A team had built a RAG system to answer internal policy and product questions from a growing document set. Early demos looked great — the model answered fluently, in the right tone, citing what looked like the right sections. Then support staff started flagging answers that were confidently wrong: correct-sounding policy language that didn't match what the actual documents said.
The root cause wasn't the language model. It was retrieval. The vector search was returning chunks that were semantically similar to the question but not actually the right source — close enough that the model treated them as ground truth and extrapolated a plausible answer on top. There was no layer checking whether retrieval had actually found the right thing before generation ran.