Chapter 11: Capstone
Time: 45 minutes. Cost: $0. The lab reuses the multi-hop dataset from Chapter 10 and the reranking model from Chapter 9, and needs no Docker. Its evaluation is the second-heaviest run in the track: about a minute on 3 to 4 CPU cores on a fast laptop, several times longer on an older one. See the "Will my computer handle it?" section of Setup.
- The capstone puts vector search, keyword search, a reranker, and links into one retrieval stack, and adds the stages one at a time to see what each is worth.
- Order matters. Running the reranker last threw away the paragraphs the links had found, and did worse than no reranker at all.
- Measure the whole stack on your own questions, inspect the failures one step at a time, and keep measuring after launch.
You have measured every piece of a retrieval system on its own. Vector search in Chapters 1 to 7. Keyword search and fusion in Chapter 8. A reranker in Chapter 9. Names and links in Chapter 10. Each one helped, on its own test.
A real system runs them together, and that raises questions no single chapter could answer. Does each piece still help when the others are already there? Does the order matter? Where does the time go? And when one question comes back wrong, which step lost the paragraph it needed?
This capstone answers those with two tools. evaluate.py measures the whole stack on 200 questions,
adding one stage at a time. The Retrieval Inspector, a small browser app, runs the same stack on
one question at a time and shows the ranking after every step.
The stack​
Every stage is code you already ran. stack.py holds all of it in about 220 lines, and it is the only
file with any search logic. The evaluation script and the inspector both call its search function.
That is a habit worth keeping in your own projects: one implementation of the pipeline, with the
measurement and the user interface built on top, so they can never disagree about what the system does.
The collection is Chapter 10's: 9,769 Wikipedia paragraphs, and questions that each need two of them. Most are two-hop questions, like the one about the director of Big Stone Gap. The rest are comparison questions, which name two things and need a paragraph about each.
What each stage adds​
PART 1: 200 questions, top 5 paragraphs kept, adding one stage at a time
two-hop (166) comparison (34)
stack both found answer in both found
vector search 0.17 0.46 0.06
+ keywords (hybrid) 0.55 0.72 0.71
+ reranker 0.61 0.76 0.97
+ names and links (full stack) 0.83 0.90 1.00
Read it from the top. Vector search alone found both needed paragraphs for 17% of two-hop questions. These questions are full of names, and Chapter 8 showed that names are where embeddings struggle. Keywords more than tripled that. The reranker added a little on two-hop questions and a lot on comparison questions, where picking the right paragraph about each named person is exactly the careful reading a cross-encoder is good at. Names and links added the most on two-hop questions, because they are the only stage that can reach a paragraph the question never describes.
Every stage earned its place on this dataset. That is a result, not a rule, and Chapter 10 explained why this dataset suits a graph so well. Chapter 8's questions each come from one paragraph, so there is no second hop to follow, and Chapter 10's comparison questions showed what links do then: take places away from search results. The point of the table is the method: add one stage at a time, on your own questions, and keep the stages that move the number you care about.
The comparison column rests on only 34 questions, so treat its exact values loosely. One question is 3 points.
The order matters​
Here is the same stack, with the same four stages, in a different order:
PART 2: the same stages, in a different order
two-hop (166) comparison (34)
stack both found answer in both found
names and links, no reranker 0.77 0.84 1.00
reranker, then names and links 0.83 0.90 1.00
names and links, then reranker 0.63 0.75 0.97
Run the reranker last, and two-hop results drop from 0.83 to 0.63. That is worse than leaving the reranker out entirely.
Think of judging puzzle pieces one at a time by asking whether each looks like the picture on the box. The piece that completes the picture often looks like nothing on its own.
The reason is in how a cross-encoder works. It reads the question and one paragraph together and asks: does this paragraph answer this question? For the paragraph about Adriana Trigiani, the honest answer is no. It never mentions Big Stone Gap. It is only relevant because of what another paragraph says, and the reranker never sees the two side by side. So it ranks her low, and she falls out of the top 5:
Names and links, then reranker:
1. Big Stone Gap (film) (name in question) <- needed
2. Nola (film) (hybrid)
3. Manhattan Romance (hybrid)
4. Just Another Romantic Wrestling Comedy (hybrid)
5. I Love NY (2015 film) (hybrid)
Nothing is broken here. Each stage did its job. The reranker is a better judge of whether one paragraph answers a question. The graph finds paragraphs that only matter in combination. Put the judge after the graph, and it throws away exactly what the graph found. Put it before, and it picks better seeds for the graph to start from.
This is the most useful thing I know about building retrieval pipelines, and it does not show up when you test each stage alone. Every stage makes an assumption. The reranker assumes relevance is a property of one paragraph. The graph assumes it can be a property of two. When you combine stages, check that a later one does not undo an earlier one.
Where the time goes​
PART 3: what each stage cost per question, full stack
vector search 0.9 ms
hybrid search 23.8 ms
reranker 159.1 ms
names and links 0.1 ms
The reranker takes about 87% of the time, as Chapter 9 predicted. The graph is almost free. Keyword search is slower than it would be in a real search engine, because this is BM25 written in plain Python, scoring every paragraph that shares a word with the question.
If I had to make this stack faster, I would start with the reranker: fewer candidates, a GPU, or a hosted reranking service. The lab's Try this section checks the other direction. Giving the reranker 50 candidates instead of 20 took its time from about 160 to about 420 milliseconds per question and did not change the results at all.
Why it runs in memory​
You might expect the capstone to run inside the Chapter 7 database. I tried that first. Vector search and the links table move into PostgreSQL easily, as Chapters 7 and 10 showed. Keyword search is the problem. PostgreSQL's built-in full-text ranking has no inverse document frequency, as Chapter 8 explained, so a rare name counts no more than a common word like "film." When I ran keyword search alone on these 1,000 questions, with Lab 10's tables loaded into the Chapter 7 database, PostgreSQL's ranking put both paragraphs of a two-hop question in the top 5 for 36% of them, against 47% for BM25.
So the capstone keeps everything in memory, where each stage is a few lines of Python you can read and
change, and the inspector starts in seconds. In production, every stage has a database home. The vectors
go in pgvector or a vector database. BM25 comes from a search engine or an extension that provides it,
such as ParadeDB's pg_search for PostgreSQL. The links are a table. The reranker runs in your
application, or as a hosted service.
After launch​
The table above is the customer check from Chapter 3: real questions with known answers, run against the whole stack. Keep the script and the questions, and rerun them whenever you change a stage, the model, or the way documents are cut into paragraphs.
The capstone uses exact vector search, so there is no index to drift. In production you would likely put the vectors behind an approximate index, and then you need the kitchen check too. Chapter 7 shows how to run it inside PostgreSQL with index scans turned off, and the trap that can make it report a perfect score that is not real.
The inspector​
The inspector opens on the Big Stone Gap question with every stage on. It shows the top 5 a language model would read, whether both needed paragraphs made it, and a table with one column per step: vector search, hybrid search, reranker, names and links. Each column is the ranking after that step, so you can follow one paragraph from left to right and see the step where it arrived or disappeared.
That table is where debugging actually happens. A number like 0.83 tells you how often the stack works. It does not tell you why it failed on the other 17%. For that, you pick a failing question and look. In the Big Stone Gap example, the Trigiani paragraph is 10th after hybrid search, gone from the top 10 after the reranker, and 4th after names and links. Turn on Run the reranker after names and links, and you can watch her disappear.
You can also type your own question. The inspector embeds it with the same model that embedded the paragraphs. Try "In which city was the band that recorded Dead but Rising formed?" The song's paragraph names the band, Volbeat, and the link brings in Volbeat's paragraph, which has the answer. Turn every stage off, and Volbeat is gone.
Hands-on lab: measure the stack, then inspect it​
You will run the evaluation, read the three tables, and then open the inspector to follow individual questions through every step.
Full instructions: download the
Vector Databases labs ZIP. If you did not do
Lab 10, first build the multi-hop dataset by following the "Multi-hop dataset" section of
labs/vector-databases/dataset/README.md. Then open labs/vector-databases/11-capstone and follow its
README.
The evaluation's quality numbers should match closely. Timings depend on your CPU, and the first run
also downloads the reranking model. The inspector runs at http://localhost:8501.
Checkpoint​
Why does running the reranker after names and links make two-hop results worse than having no reranker at all?
The reranker scores each paragraph alone against the question. A second-hop paragraph, such as the one about Adriana Trigiani, does not answer the question on its own, so it scores low and falls out of the top 5. Running the reranker first lets it choose better seeds and leaves the linked paragraphs alone.
The full stack finds both paragraphs for 83% of two-hop questions. How would you find out why it fails on the rest?
Pick failing questions and follow them through every step, which is what the inspector's table is for. Find the step where a needed paragraph disappears, or notice that it never appeared at all, then change that step and measure again on all the questions.
Why does the capstone keep BM25 in Python instead of using PostgreSQL's full-text search?
Two-hop questions are full of rare names, and the name is usually what finds the right paragraph. PostgreSQL's built-in ranking has no inverse document frequency, so a rare name counts no more than "film" or "born." BM25 weighs the rare name heavily. In production you would get BM25 from a search engine or an extension, rather than writing it yourself.
Check Your Knowledge​
Click to start the quiz
What's next​
That is the whole track. You started with a 1960s idea, putting documents and questions in the same space, and finished with a retrieval stack you can measure and take apart one step at a time. Along the way you measured recall against exact answers, chose between an extension and a dedicated database, and found out where vectors are weak and what fixes them.
The most useful next step is to point this method at your own documents. Write 50 real questions with the paragraphs that answer them. Start with vector search, add one stage at a time, and keep only the stages that move the number. Then open a failing question and look at every step.
If you want to build on top of retrieval, Advanced Chapter 2 uses a language model to improve it: rewriting the question, HyDE, multi-hop retrieval that searches again, and retrieval that checks its own results. The Advanced capstone then evaluates a whole agent, not only the paragraphs it reads.