Re-ranking: a second, more careful look
Re-ranking is a second pass over search results. Search picks a handful of candidate passages quickly; a re-ranker then reads each candidate together with the question and puts them in a better order. Only the top few after re-ranking go into the prompt. It is one of the most dependable ways to improve a RAG system that already finds the right passage but does not rank it first.
Why search alone gets the order wrong
Vector search turns the question into one embedding and every passage into another, each on its own, and compares the two. That is what makes it fast: the passage vectors are computed once and stored. The price is that the model never sees the question and the passage side by side, so it cannot notice that a passage mentions the right subject but answers a different question about it.
Ask "How long is the warranty on brakes?" of a bike handbook. A passage about returning brakes and a passage about the warranty on frames both look close. The passage that actually says "brakes are covered for two years" may come third.
What a re-ranker does differently
A re-ranker takes the question and one passage as a single input and returns one number: how well this passage answers this question. Because it reads both together, it can weigh every word of the question against every word of the passage. This kind of model is called a cross-encoder.
The cost is speed. Nothing can be computed in advance, so scoring 10,000 passages would mean 10,000 model runs per question. That is why re-ranking is always the second stage and never the first.
The two-stage pattern
- Retrieve widely. Use fast search, ideally hybrid search, to fetch more candidates than you need: 10 to 50 is common.
- Re-rank. Score each candidate against the question with the slower, more careful model.
- Keep the best. Put only the top three to five into the prompt.
The first stage is tuned for recall: do not miss the right passage. The second is tuned for precision: put it first.
Two kinds of re-ranker
- Cross-encoder models. Small models trained for exactly this job. They are fast enough to score a few dozen passages in well under a second on a server, cheap to run, and consistent. Most production systems use one.
- A general language model. You can ask a chat model to score each passage from 0 to 10. This needs no extra model and can follow instructions such as "prefer the most recent policy", but it is slower, costs tokens for every candidate on every question, and its scores can vary between runs.
How many candidates?
- Too few and the re-ranker has nothing to rescue: if the right passage was 12th and you re-rank 10, it stays lost.
- Too many and each question gets slower and, with a language model, more expensive, for little gain.
- A sensible start is to re-rank three to five times as many passages as you will keep, then measure with real questions.
When it helps and when it does not
Re-ranking helps when the right passage is among the candidates but not at the top. It also lets you put fewer passages in the prompt, which lowers the token cost of every answer.
It does not help when search never retrieves the passage, when the answer is split across two passages, or when the document does not contain the answer. Fix chunking and retrieval first.
Trying it
In the RAG tool, open Settings and set "Re-rank" to the small model in your browser (a one-time download) or to your own model if you have connected one. "Re-rank candidates" sets how many passages search hands over. Each retrieved passage then shows how far it moved, and the Step by step panel lists every candidate with its position and score before and after.
The browser model is a small cross-encoder trained on English web search questions. It shows the idea well, but a production re-ranker would be larger and may be trained for your language or subject.
More guides
What is RAG? · How to choose a chunk size for RAG · What are embeddings? · What are tokens, and what do they cost? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes · What is GraphRAG?