Writing · 4 min read

Why RAG retrieval fails

Most bad RAG answers are retrieval failures in disguise. A field guide to the common failure modes, how to recognise each one, and what usually fixes it.

When a retrieval-augmented generation (RAG) system gives a wrong answer, the instinct is to blame the language model and rewrite the prompt. In practice the model is often doing exactly what it was asked: answering from the passages it was given. The passages were wrong.

So the first diagnostic step is always the same. Look at what was retrieved before looking at what was generated. If the right passage is not in the context, no prompt will fix the answer.

Below are the failure modes I check for, roughly in the order they occur in the pipeline.

1. The text was never extracted correctly

Retrieval cannot find what parsing destroyed. Typical causes:

  • scanned PDFs indexed without OCR, so they contain no text at all;
  • multi-column layouts read across the columns, interleaving unrelated sentences;
  • tables flattened into a stream of numbers with no headers;
  • headers, footers and page numbers repeated into every chunk;
  • documents in another language, or with broken encoding.

How to recognise it: open the stored text for a document you know the answer is in, and read it. This takes five minutes and is skipped surprisingly often.

2. The answer was split across chunk boundaries

Fixed-size chunking cuts wherever the token counter says so: between a heading and its paragraph, between a table and its caption, between a question and its answer in an FAQ. The retriever then finds half the evidence, or a chunk that scores well but lacks the key sentence.

What usually helps: chunk on document structure (sections, paragraphs, table boundaries), keep the section title attached to each chunk, and use modest overlap. Bigger chunks are not automatically better. They dilute the similarity signal and spend context budget.

3. The question and the document use different words

A user asks about “staff holidays”; the policy says “annual leave entitlement”. Pure keyword search (BM25) misses this. Dense embeddings were designed for this case, and it is the main reason they exist.

4. The question contains the exact words, and dense search ignores them

The reverse failure is just as common and less discussed. Part numbers, error codes, names, legal article references, rare technical terms: embedding models compress text into a fixed-size vector, and exact identifiers often do not survive that compression well. “Error E-4071” retrieves passages about errors in general.

What usually helps: hybrid retrieval, meaning lexical and dense search run in parallel with the results fused (reciprocal rank fusion is a simple, robust default), so each method covers the other’s blind spot. Whether hybrid actually helps on your data is an empirical question; see the LAB/001 measurements when they are published.

5. The right passage was retrieved, but ranked too low

Retrievers are tuned for recall: getting the relevant passage somewhere in the top 50 or 100. The language model only sees the top few. If the right passage sits at rank 23 and the context holds 8, it might as well not exist.

How to recognise it: measure recall at several depths (top 5, 20, 100). If recall@100 is high but recall@5 is low, the retriever is finding the evidence and the ranking is losing it. A cross-encoder reranker over the top candidates is the standard fix; it costs latency, which you should measure too.

6. Filters and metadata were ignored

“What did the 2024 contract with supplier X say about penalties?” contains two hard constraints, a year and a supplier, that semantic similarity will happily ignore, returning the most similar penalty clause from any contract. Constraints that the user states explicitly should become metadata filters, not hopes.

7. The index is stale or incomplete

Documents updated last week, permission changes, deleted files that are still retrievable. These are ordinary data-engineering problems: incremental indexing, deletion handling and a way to see what the index currently contains.

8. The question needs more than one passage

“Which of our three offices had the highest turnover growth?” requires several facts from different documents and a comparison. Single-shot top-k retrieval often returns three passages about one office. Such questions need decomposition (retrieve per sub-question) or a structured data source, and sometimes the honest answer is that a document search is the wrong tool for an analytical question.

9. The answer does not exist

Some questions have no answer in the corpus. A good system says so. A system without a “not found” path will generate something plausible from the nearest passages, and that looks exactly like failure mode 5.

Measure retrieval separately from generation

All of the above leads to one practice: evaluate retrieval on its own. Build a set of real questions, record which passages contain each answer, and measure recall and ranking quality (for example recall@k and nDCG@10) for every change to parsing, chunking, models or fusion. Generation quality can then be judged on top of retrieval you know is working.

Without that separation, every change is a guess, and RAG systems tend to get tuned by anecdote, one complaint at a time.