Seven common RAG mistakes
Most disappointing RAG systems fail for a small number of ordinary reasons. None of them needs a bigger model to fix. Here are seven, roughly in the order they are worth checking.
1. Never looking at what was retrieved
The most common mistake is judging only the final answer. When an answer is wrong, the first question should be: was the right passage in the prompt? If it was not, no model could have answered correctly, and changing the model or the prompt wording will not help.
Fix: always look at the retrieved passages. The RAG tool shows them, with their scores, before anything is sent to a model.
2. Passages cut in the wrong places
Splitting every fixed number of characters is easy to code and often cuts a sentence in half, separates a heading from its section, or pulls a label away from its value. The answer then exists in the document but in no single passage.
Fix: split at sentence or heading boundaries and add a little overlap. Look at the result in the chunking visualizer; the count of passages that end mid-sentence should be low.
3. One chunk size for everything
A size that suits a policy document is rarely right for a table-heavy manual or a set of short questions and answers. Sizes copied from a tutorial are tuned for the tutorial's documents, not yours.
Fix: choose the size per kind of document and measure. The chunk size guide gives starting points.
4. Relying on meaning search alone
Embedding models are good at paraphrase and weak at exact strings. Searches for an order number, an error code or an unusual name can return passages that are "about the same sort of thing" but not the one that contains the string.
Fix: combine word matching with meaning matching. See hybrid search.
5. Retrieving too much, or too little
Retrieving one passage leaves no margin if the best match is slightly off. Retrieving twenty buries the useful passage among unrelated ones, raises the cost of every question and can lead the model onto side topics. Padding the prompt with weak matches is a frequent cause of answers that are confidently about the wrong thing.
Fix: start with three to five passages, and set a minimum score so that weak matches are left out even if that means fewer passages. Leave out near-duplicates so the slots are not all taken by copies of the same text.
6. No instruction for "the answer is not here"
If the prompt only says "answer the question", a model will try to answer even when the retrieved passages do not contain the information, filling the gap from its general knowledge. That is how invented answers appear in a system that was meant to prevent them.
Fix: tell the model to use only the provided context and to say it does not know when the context is insufficient. Ask it to quote or point to the passage it used, so a person can check.
7. Tuning by feel
Changing a setting, trying two questions and deciding "that seems better" is not a test. The two questions may have improved while ten others got worse.
Fix: keep a small set of test questions with known answers and re-run it after every change. The method is described in how to test RAG retrieval, and the Tune tab of the RAG tool runs it for you.
Three smaller ones
- Unreadable source text. A scanned PDF is a picture of text, not text. Nothing can be retrieved from it until it has been through character recognition. Check that copied text from the file looks right before blaming retrieval.
- Stale documents. If a document is updated but its passages are not re-embedded, the system keeps answering from the old version.
- Ignoring cost. Large passages and a high number retrieved multiply the tokens in every prompt. Estimate it early with the token counter.
A quick diagnostic
- Is the answer actually in the document, in readable text?
- Is it inside a single passage, or split across two?
- Is that passage among those retrieved? At what position?
- If yes, does the prompt tell the model to answer only from the context?
The first "no" tells you where to work. You can walk through all four on your own document in the RAG tool.
More guides
What is RAG? · How to choose a chunk size for RAG · What are embeddings? · What are tokens, and what do they cost? · Hybrid search: words plus meaning · How to test RAG retrieval