Chapter 9: Rerankers
Time: 35 minutes. Cost: $0. The lab downloads an 80 MB reranking model the first time and runs it on your CPU. It reuses the shared dataset and does not need Docker. It is the heaviest lab in the track: about 70 seconds on 4 CPU cores on a fast laptop, and several times longer on an older one. Read the "Will my computer handle it?" section of Setup first.
- A reranker is a slower, more careful model that rereads the top 20 or so search results together with the question and puts them in a better order.
- In the lab it moved the right paragraph into first place for 83% of questions, up from 71%. It cannot find a paragraph the search missed, so fix the search first.
- It is the most expensive step in the search: about 176 milliseconds on a laptop, against about 1 for the search itself. Rerank fewer candidates to trade a little quality for a lot of speed.
Chapter 8 ended on a gap. The best hybrid search put the right paragraph in its top 10 for 92% of questions, but in first place for only 70%. Most of the time the answer was there. It just was not on top.
That gap matters more than it looks. A RAG application rarely sends 10 paragraphs to the language model. It sends three, or one. If the right paragraph is sitting at rank 6, the model never sees it, and it answers confidently from whatever was ranked first.
Think about how a good research librarian works. First they pull 20 likely books off the shelves using the catalog. That part is fast and a little rough. Then they read the first page of each one with your question in mind and hand you the best three. That second step is a reranker: a slower, more careful reader that only looks at what the fast search already found.
Two ways to compare a question and a paragraph​
Every search in this track so far has used a bi-encoder. The question and the paragraph are embedded separately, each into its own vector, and compared with a dot product. The paragraph's vector was computed before anyone asked anything. That is the whole reason vector search is fast: the hard work happens ahead of time, and search is just math on stored vectors.
It is also the weakness. A paragraph's vector has to stand for every question that paragraph could ever answer. It cannot know that this particular question hinges on the word "Italian." It is like judging job applicants only by résumés they wrote before they knew which job they were applying for.
A cross-encoder is the interview. It reads the question and the paragraph together, in one pass through the model, and outputs a single relevance score. Every word of the question can look at every word of the paragraph. It is far more precise. It also means there is nothing to compute ahead of time. Every new question needs a fresh pass for every paragraph it is compared against.
The Sentence-BERT paper from Chapter 1 put a number on that cost. Finding the most similar pair among 10,000 sentences by running BERT on every pair takes about 65 hours. With separately embedded sentences, about 5 seconds, with similar accuracy. That is why a cross-encoder never searches the whole collection. It only reranks a short list.
Where rerankers came from​
The idea took off when researchers trained BERT, an early transformer model, to read a search query and a passage together and score the pair, then used it to reorder a regular search engine's results. They tested it on MS MARCO, a large public set of real Bing search queries with the passages people marked as relevant, and went straight to the top of its leaderboard.
That two-stage pattern, fast retrieval and then a careful reranker, has been standard ever since. The model in this lab comes straight from that line: a small 6-layer MiniLM cross-encoder trained on MS MARCO passage ranking, published by the Sentence-Transformers project. The lab runs it through fastembed, a library that runs models in ONNX format on your CPU, so it installs on Windows, macOS, and Linux without PyTorch.
Where this came from: Nogueira and Cho, and MS MARCO
Microsoft released MS MARCO in 2016. In January 2019, Rodrigo Nogueira and Kyunghyun Cho fine-tuned BERT to score a query and a passage together and used it to rerank search engine results on MS MARCO. Their system went to the top of the MS MARCO passage ranking leaderboard and beat the previous best by 27% on its main metric.
ColBERT, in the next section, was published by Omar Khattab and Matei Zaharia at SIGIR in 2020. The study of language models as rerankers is by Weiwei Sun and colleagues, from 2023.
What reranking fixes​
The lab takes the first 200 questions, gets the top 20 candidates from plain vector search and from Chapter 8's hybrid search, and reranks the top 5, 10, or 20 with the cross-encoder. A cross-encoder is slow, which is why it uses 200 questions instead of all 1,000. On these 200, hybrid search puts the right paragraph first 71% of the time, close to the 70% Chapter 8 measured on all of them:
PART 2: rerank the top N with the cross-encoder
first stage reranked found@1 found@10 ceiling
vector none 0.41 0.57
5 0.51 0.57 0.54
10 0.54 0.57 0.57
20 0.59 0.64 0.64
hybrid (Chapter 8) none 0.71 0.90
5 0.79 0.90 0.82
10 0.81 0.90 0.90
20 0.83 0.95 0.95
Reranking hybrid's top 20 moved the right paragraph into first place for 83% of questions, up from 71%. On plain vector search, first place went from 41% to 59%. That is a large gain for a model that never saw this dataset. It was trained on web search queries, and these are questions about Wikipedia.
Here is one question from the lab, before and after:
PART 4: "What are Italian dialects termed in the Italian language?"
The source paragraph is 4101, from 'Dialect'.
Hybrid search:
1. 4081 'Dialect': The other usage refers to a language that is socially subord...
2. 4101 'Dialect': Italy is home to a vast array of native regional minority la... <- source
3. 10993 'Switzerland': Aside from the official forms of their respective languages,...
After reranking the top 20:
1. 4101 'Dialect': Italy is home to a vast array of native regional minority la... <- source
2. 4106 'Dialect': Italians in different regions today may also speak regional ...
3. 4081 'Dialect': The other usage refers to a language that is socially subord...
Hybrid search found two paragraphs from the right article and ranked the general one first. The reranker read both with the question and saw that only one is about Italy.
What reranking cannot fix​
Now look at the column called ceiling. It is how often the source paragraph was anywhere in the candidates the reranker received. A reranker can only reorder what it is given. It cannot find a paragraph the first stage missed.
On vector search, reranking 20 candidates brought found@10 to 0.64, exactly the ceiling. Every question the reranker could fix, it fixed. The other 36% were never in the list. A bigger, smarter reranker would not change that number at all.
A reranker is a better judge, not a better searcher. If your first stage has poor recall, reranking gives you a well-ordered list of the wrong paragraphs. Fix recall first, which is what Chapter 8's hybrid search did, then rerank. On hybrid search, the ceiling at 20 candidates was 0.95, and that is where the real gains came from.
What it costs​
PART 3: what reranking costs on this machine's CPU
Average over 6,540 question-paragraph pairs: 9.4 ms per pair
reranked ms per question (measured on 20 questions)
5 30
10 78
20 176
For comparison, exact vector search over all 18,896 paragraphs: 1.3 ms per question
Reranking 20 candidates took 176 milliseconds per question on a recent Apple laptop, using 4 CPU threads. Exact vector search over the whole collection took about 1 millisecond. The reranker is easily the most expensive step in this retrieval stack, and on an older laptop, expect it to be several times slower.
The cost grows faster than the number of candidates. The model processes candidates in batches, and every paragraph in a batch is padded to the length of the longest one. More candidates mean a better chance of one long paragraph slowing down the whole batch.
That table is also your tuning knob. Reranking the top 10 instead of 20 got found@1 to 0.81 instead of 0.83, at less than half the cost. If I had a tight latency budget, that is where I would start, then measure whether the extra 2 points are worth about 100 milliseconds for my users.
A GPU changes the picture. The model's documentation reports about 1,800 documents per second on an NVIDIA V100. Hosted reranking APIs exist for the same reason: they run the model on hardware you do not have to manage.
Late interaction: a middle ground​
A bi-encoder squeezes a paragraph into one vector ahead of time. A cross-encoder does everything at query time. Late interaction sits in between.
Back to the librarian. Instead of reading the first page of every book while you wait, suppose they had written a small note card for every key word in every book when it arrived. At question time, they only match your words against the cards. Most of the reading was done in advance.
ColBERT works that way. It keeps one vector for every token of the paragraph instead of one for the whole paragraph, all computed ahead of time. At query time, each word of the question finds its best-matching token in the paragraph, and those best matches are added up. Words still meet words, as in a cross-encoder, but the paragraph side was done in advance. The paper reported rankings competitive with BERT-based rerankers while running two orders of magnitude faster.
The price is storage. These paragraphs average about 117 words, so each one becomes well over a hundred vectors instead of one. Later versions of ColBERT compress those vectors heavily to make that manageable. Late interaction can rerank, like the cross-encoder here, or search the whole collection with an index built for it.
Rerankers made from language models​
A language model can rerank too. Show it the question and the candidate paragraphs, and ask it to put them in order. A 2023 study tested exactly that with ChatGPT and GPT-4, sliding a window over the candidate list to handle more passages than fit in one prompt. With the right instructions, the models were competitive with, and sometimes better than, rerankers trained for the task. The authors also distilled that ranking skill into a much smaller model.
Intermediate Chapter 3 used a simple version of this: the model picked the best of hybrid search's top three. It works, and it can reason about a question in a way a small cross-encoder cannot. But every rerank is now a language model call, with every candidate paragraph in the prompt. For 20 paragraphs like these, that is a few thousand tokens per question.
My default is a cross-encoder. It is cheap, predictable, and runs anywhere. I would consider a language model as the reranker when queries are few and each one is valuable, or when choosing the right passage takes real reasoning about the question.
One more distinction, because the two get mixed up. Advanced Chapter 2 covers HyDE, which rewrites what you search for before the first stage. That can raise the ceiling. A reranker works after the first stage and only reorders. They solve different problems and work well together.
Hands-on lab: rerank vector and hybrid results​
You will build the two first-stage searches from Chapter 8, rerank their top candidates with a cross-encoder, measure found@1 and found@10 against the ceiling, time the reranker, and look at one question before and after.
Full instructions: download the
Vector Databases labs ZIP. If you have not
built the shared dataset yet, follow labs/vector-databases/dataset/README.md first. Then open
labs/vector-databases/09-rerankers and follow its README.
The output is the tables in this chapter. Quality numbers should match closely. Timings depend heavily on your CPU, and the first run also downloads the model.
Checkpoint​
Why not skip the first stage and run the cross-encoder over all 18,896 paragraphs?
A cross-encoder needs a fresh pass for every question and paragraph pair, and nothing can be computed ahead of time. At about 9 ms per pair on the lab's CPU, scoring all 18,896 paragraphs would take close to three minutes per question. A first stage narrows the list to 20 in about a millisecond.
Reranking vector search's top 20 brought found@10 to 0.64 and no higher. Why, and what would you fix first?
The source paragraph was in those 20 candidates only 64% of the time, and a reranker can only reorder what it receives. Every fixable question was already fixed. Improve the first stage's recall, for example with hybrid search, whose top 20 contained the source 95% of the time.
What does late interaction compute ahead of time that a cross-encoder cannot, and what does it cost?
It stores one vector per token of every paragraph, computed before any question arrives. At query time, it only matches question words against those stored token vectors. The cost is storage: each paragraph becomes dozens to hundreds of vectors instead of one.
Check Your Knowledge​
Click to start the quiz
What's next​
Everything up to now has assumed the answer lives in one paragraph. Some questions do not work that way. "Which company acquired the startup founded by the author of this paper?" needs one fact to find the next. No amount of reranking helps when the second paragraph shares no words or meaning with the question. Chapter 10 stores a graph of entities and relationships next to the vectors and follows the links.