How to evaluate a RAG system without fooling yourself
A small, honest evaluation set beats a large, vague one. How to build one, which numbers to track, and the mistakes that make RAG systems look better than they are.
Most RAG systems are evaluated by their builders asking a dozen questions they already know the answers to, and nodding. Then the system meets real users, and nobody can say whether last week’s change made it better or worse.
An evaluation does not need to be large to be useful. It needs to be honest. This is the process I follow.
1. Start from real questions
Collect questions the way users actually ask them: from support tickets, search logs, emails, or by asking the people who will use the system to write down what they would ask. Questions written by the system’s builder are systematically easier. They use the documents’ own vocabulary, because the builder has read the documents.
Fifty to a hundred real questions is a good start. Include questions that have no answer in the documents. A system that never says “I don’t know” will otherwise look perfect.
2. Label the evidence, not just the answer
For each question, record which document passages contain the answer. This is the expensive part, and the part that makes everything else possible: with passage labels you can measure retrieval directly instead of inferring it from the final answer.
Labelling tips:
- Label with the people who know the domain, not the developer.
- Record all passages that answer the question, not just the first one found.
- Keep the labels with the documents’ versions. When a document changes, the label may become wrong.
3. Measure retrieval and generation separately
They fail for different reasons and are fixed in different places.
Retrieval (did the right passages reach the model?):
- Recall@k: the share of labelled passages found in the top k, for the k you actually pass to the model and for a deeper k like 100. A large gap between the two means ranking is the problem, not recall.
- nDCG@10 or MRR: whether the relevant passages are near the top.
Generation (given the passages, was the answer right?):
- Correctness against a reference answer.
- Groundedness: is every claim supported by the cited passage?
- Abstention: does it say “not found” on the unanswerable questions, and only on those?
When an answer is wrong, the labels tell you which half failed. If the passage was not retrieved, prompt changes are wasted effort.
4. Treat LLM-as-judge as an instrument that needs calibration
Using a language model to grade answers scales well and is often the only practical option for generation quality. It is also a measuring instrument with its own biases: it favours longer answers, answers in its own style, and confident tone.
Before trusting it, grade 30–50 answers by hand and check how often the judge agrees with you. Keep that hand-graded set and re-check whenever the judge model or its prompt changes.
5. Don’t tune on your test set
Every time you look at a failing question and adjust chunking, prompts or retrieval to fix it, that question stops being a fair test. Keep a development set to iterate on and a held-out set you only run before decisions. With a small total, even a 70/30 split is better than none.
6. Compare per question, not just on averages
A change that moves the average nDCG@10 from 0.41 to 0.42 on 60 questions is usually noise. Look at the paired view instead: on how many questions did the change help, hurt, or make no difference? A change that helps 9 questions and hurts 8 is not an improvement, whatever the average says. When it matters, a paired bootstrap or sign test over the questions gives a defensible answer.
In the LAB/001 retrieval measurements, averages and per-query counts sometimes tell different stories, which is exactly why both are reported.
7. Report cost and latency next to quality
A reranker that adds two points of nDCG@10 and 200 milliseconds per query is a trade-off, not a free improvement. Put quality, latency (median and 95th percentile, measured one query at a time) and cost per query in the same table, so decisions are made on all three.
8. Make it a regression test
The evaluation is only useful if it runs on every change: new documents, a new chunker, a new model version from the vendor. Keep the questions, labels and the script in the repository, run them in CI or before each release, and record the results with the date and configuration.
The minimum viable version
If this all sounds like a lot: 50 real questions, the passages that answer them, recall@k and a hand check of 20 answers. That takes a few days, and it is the difference between tuning a system and guessing.