Re-ranking: a second, more careful look

Re-ranking is a second pass over search results. Search picks a handful of candidate passages quickly; a re-ranker then reads each candidate together with the question and puts them in a better order. Only the top few after re-ranking go into the prompt. It is one of the most dependable ways to improve a RAG system that already finds the right passage but does not rank it first.

Why search alone gets the order wrong

Vector search turns the question into one embedding and every passage into another, each on its own, and compares the two. That is what makes it fast: the passage vectors are computed once and stored. The price is that the model never sees the question and the passage side by side, so it cannot notice that a passage mentions the right subject but answers a different question about it.

Ask "How long is the warranty on brakes?" of a bike handbook. A passage about returning brakes and a passage about the warranty on frames both look close. The passage that actually says "brakes are covered for two years" may come third.

What a re-ranker does differently

A re-ranker takes the question and one passage as a single input and returns one number: how well this passage answers this question. Because it reads both together, it can weigh every word of the question against every word of the passage. This kind of model is called a cross-encoder.

The cost is speed. Nothing can be computed in advance, so scoring 10,000 passages would mean 10,000 model runs per question. That is why re-ranking is always the second stage and never the first.

The two-stage pattern

  1. Retrieve widely. Use fast search, ideally hybrid search, to fetch more candidates than you need: 10 to 50 is common.
  2. Re-rank. Score each candidate against the question with the slower, more careful model.
  3. Keep the best. Put only the top three to five into the prompt.

The first stage is tuned for recall: do not miss the right passage. The second is tuned for precision: put it first.

Two kinds of re-ranker

How many candidates?

When it helps and when it does not

Re-ranking helps when the right passage is among the candidates but not at the top. It also lets you put fewer passages in the prompt, which lowers the token cost of every answer.

It does not help when search never retrieves the passage, when the answer is split across two passages, or when the document does not contain the answer. Fix chunking and retrieval first.

Trying it

In the RAG tool, open Settings and set "Re-rank" to the small model in your browser (a one-time download) or to your own model if you have connected one. "Re-rank candidates" sets how many passages search hands over. Each retrieved passage then shows how far it moved, and the Step by step panel lists every candidate with its position and score before and after.

The browser model is a small cross-encoder trained on English web search questions. It shows the idea well, but a production re-ranker would be larger and may be trained for your language or subject.

More guides

What is RAG? · How to choose a chunk size for RAG · What are embeddings? · What are tokens, and what do they cost? · Hybrid search: words plus meaning · How to test RAG retrieval · Seven common RAG mistakes · What is GraphRAG?