How to test RAG retrieval
When a RAG answer is wrong, there are two possible causes: the right passage was never retrieved, or it was retrieved and the model still answered badly. The first is far more common, and it can be measured without calling a model at all. This guide describes a simple method.
Why test retrieval separately
Judging final answers is slow and subjective, and it mixes two problems together. Retrieval, by contrast, has a clear right answer: for a given question, either the passage that contains the answer is in the retrieved set or it is not. That makes it quick to check, cheap to repeat, and fair to compare between settings.
Step 1: write test questions
Collect five to twenty questions that people would really ask about the document. For each one, find the place in the document that answers it and note a short word or phrase that only the right passage contains. This is the "must contain" text.
- Question: "How many days do I have to return a bike?" Must contain: "30 days".
- Question: "How long is the frame warranty?" Must contain: "five year".
Good test sets include a mix:
- Direct questions that reuse the document's own words.
- Rephrased questions that mean the same but use different words, as real users do.
- Exact-term questions about a code, name or number.
- Questions from different parts of the document, not only the first page.
Choose the "must contain" text carefully. It should be specific enough that only the right passage has it, and short enough that a different split does not cut it in two.
Step 2: measure
For each question, run retrieval and look at the passages returned. Two numbers summarise the result:
- Found: in how many questions was the right passage among those retrieved? If you retrieve three passages, this is often written "recall at 3".
- Position: when found, was it first, second or third? Higher is better, because models pay most attention to what comes first and because a passage found at position one will survive if you later retrieve fewer.
On the Tune tab of the RAG tool, "Test current settings" shows both for every question.
Step 3: change one thing at a time
Now vary the settings and measure again. The ones with the most effect are usually:
- passage size and overlap,
- the splitting method (fixed, by sentence, by heading),
- how many passages are retrieved,
- matching by words, by meaning, or a blend of both.
Change one and keep the rest fixed, or you will not know which change made the difference. "Try many settings" on the Tune tab runs fifteen combinations of method and size in one go and sorts them best first.
Step 4: read the result sensibly
- Prefer the simplest setting among the best. If several settings find every answer, choose the one with fewer or smaller passages; it will cost less per question.
- Look at the failures. One question that fails under every setting usually points to a real gap: the answer is spread across sections, or the document does not actually say it.
- Do not over-fit. With ten questions, the difference between 9 and 10 found may be luck. Add more questions before trusting a small gap.
- Re-test when documents change. A setting tuned on one manual may not suit the next.
After retrieval: checking the answer
Once retrieval is reliable, the remaining checks are about the model's answer. Two questions are worth asking of each one: is every statement in the answer supported by the retrieved passages, and does the answer actually address the question? These need a person or a second model to judge, which is why it pays to get retrieval right first.
A ten-minute routine
- Load the document into the RAG tool.
- Open the Tune tab and enter five questions with their "must contain" text.
- Press "Try many settings".
- Press "Use" on the best row, go back to Ask, and try a few questions by hand.
That is enough to replace guessing with a measured choice.
More guides
What is RAG? · How to choose a chunk size for RAG · What are embeddings? · What are tokens, and what do they cost? · Hybrid search: words plus meaning · Seven common RAG mistakes