"The model hallucinated" is usually the wrong diagnosis for a RAG failure. The model is often doing exactly what it's supposed to do — generating a fluent, confident answer from the context it was given. The failure is almost always upstream of generation:
- Wrong chunks retrieved. The retriever pulled content that's topically related but doesn't actually answer the query, and the model answers from it anyway rather than recognizing the mismatch.
- Chunking that lost meaning. A document split at a fixed character count can cut a sentence, a table, or a conditional clause in half — the retrieved chunk reads coherently but no longer means what the source document meant.
- No groundedness check. Nothing in the pipeline verifies that the generated answer is actually supported by the retrieved text, so a model that extrapolates slightly beyond the source passes through unflagged.
That's why "confident" is the operative word — a RAG system with these gaps doesn't fail visibly. It produces an answer that reads exactly like a correct one.