Vector Databases Overview
Loading progress saved in this browser...
- A vector database finds the stored text closest in meaning to a question. This track opens it up and shows how it does that, one idea per chapter.
- You need no computer science background, and no math beyond arithmetic. Each chapter starts with the idea in plain words. The deeper history and math sit in sections you can open if you want them.
- Every chapter has a lab you run on your own laptop, so you see each idea work, and fail, on real data.
Before you start: this track assumes you have read Foundations Chapter 4: What Is an Embedding? and Chapter 5: What Is a Vector Database?. You do not need Intermediate or Advanced.
New here? Start with this track's Setup page. It covers everything the labs need, including Docker, which you will use from Chapter 7.
Where you are picking up​
In Foundations you turned sentences into embeddings, stored them in a vector database, and got back the closest matches. It worked, and you did not need to know how. The database was a black box: text in, nearest neighbors out.
That is the right way to start. It stops being enough the first time something goes wrong in production. Search returns a confident wrong answer. Results get worse after you add a million documents. A vendor asks you to choose between two index types you have never heard of, HNSW and IVF, or between two ways of measuring similarity, cosine and dot product. Your team asks whether you need a dedicated vector database at all, or whether the PostgreSQL database you already run will do. Answering any of those means opening the box.
Do you still need vector search?​
It is a fair question in 2026. The newest models read a million tokens in one request. Google's documentation puts that at about eight novels. Why search for the right paragraphs when you could hand the model all of them?
Sometimes you should. Anthropic's 2024 guidance on retrieval says that if your knowledge base is under about 200,000 tokens, roughly 500 pages, you can put the whole thing in the prompt. Prompt caching then makes sending the same large prompt again much cheaper. For a team handbook or a single product manual, I would start there and skip the vector database.
Past that size, four things bring you back to search:
- Size. Most real collections do not fit. The Wikipedia paragraphs this track uses from Chapter 4 come to about 2.9 million tokens, nearly three times a million-token window. A company's shared drive is far bigger.
- Cost per question. You pay for every token the model reads, on every question. At October 2026 prices, a full million-token prompt to Claude Sonnet 5.5 costs about $2 in input alone, or about $0.20 when the same prompt is cached. Retrieving the five paragraphs that matter, about a thousand tokens, costs a fifth of a cent. At a few thousand questions a day, that difference decides whether the product is affordable.
- Accuracy. A longer prompt is not a better-read prompt. The 2023 "Lost in the Middle" study found that models answered best when the passage they needed sat near the start or end of a long input, and worse when it sat in the middle. The NoLiMa benchmark (2025) removed the shared words between question and answer, the same vocabulary problem Chapter 1 opens with. At 32,000 tokens, 11 of the 13 models it tested scored below half of what they managed on short inputs. Google's own long-context guide warns that accuracy drops when the model has to find several facts at once. Newer models do better than the ones in those studies. I would still test before trusting a model to read page 4,000 as carefully as page 1.
- Control. Search returns specific documents. You can show them as sources, leave out the ones a user is not allowed to see, and update one document without rebuilding a giant prompt.
Not every system that finds text uses vectors, though. Coding tools are the clearest split. Claude
Code's creator wrote in February 2026 that early
versions used RAG (search first, then hand the results to the model) with a local vector database,
and that the team found agentic search, the model running plain text searches like grep itself,
generally worked better. Cursor reported the other side in
November 2025: adding semantic search, which searches by meaning with vectors, on top of grep
raised its agent's accuracy by 12.5% on average, and the combination worked best. Code is full of
exact names, which is where keyword search is strong. Chapter 8 measures the same pattern on ordinary
text.
My read: long context changed when you need retrieval, not whether. Small, stable knowledge goes in the prompt. Large, changing, or private knowledge needs search, and good search usually mixes vectors with keywords. Vector search also has jobs that have nothing to do with chatbots: recommendations, finding duplicates, and searching images. Those are not going away.
What is different now​
This track opens it, starting from the beginning. Vector search was not invented for chatbots. The idea of putting documents and questions in the same space and measuring the distance between them dates back to information retrieval research in the 1960s. The algorithm that finds the nearest one fast is a separate story that runs from a 1951 Air Force report to the graph indexes most vector databases offer today.
Each chapter takes one step along that line, explains why the step was needed, and has you build or measure it yourself. Chapters 1 to 3 use 16 help-desk articles, small enough to see every score. From Chapter 4, where speed starts to matter, the labs share a bigger collection: 18,896 Wikipedia paragraphs and 1,000 questions about them. Chapters 10 and 11 switch to 1,000 questions that each need two paragraphs to answer.
What you will be able to do by the end​
By the end of this track, you will be able to explain how text becomes a vector and why "close" means "similar." You will be able to choose a distance metric deliberately, and say when exact search is good enough and when it is not. You will be able to tune HNSW, the most common index, by checking its answers against exact search, and estimate how much memory a billion vectors need. You will be able to decide between a vector extension on a database you already run and a purpose-built vector database. And you will build a retrieval stack that combines keyword search, vector search, reranking, and a knowledge graph, then measure whether each piece earns its place.
Chapters ahead​
- Setup: uv, an embedding model, and Docker on macOS, Windows, or Linux.
- Chapter 1: Meaning as Geometry. How text became vectors, from word counts to neural embeddings. Lab: search the same articles with TF-IDF, LSA, and an embedding model.
- Chapter 2: The Math of Similarity. Dot product, cosine, and Euclidean distance, which one fits which job, why they agree on normalized vectors, the curse of dimensionality, and what similarity scores really look like. Lab: measure all of it on real embeddings.
- Chapter 3: Exact Search. The nearest-neighbor rule, brute force, why tree indexes break in high dimensions, and the two kinds of ground truth you check any search against. Lab: time brute force and build a k-d tree.
- Chapter 4: Approximate Search. Locality-sensitive hashing and inverted file indexes: trading a little accuracy for a lot of speed, and recall, the number that tells you how much you traded. Lab: build both by hand and measure recall against exact answers.
- Chapter 5: HNSW. The graph index most vector databases use today, from the small-world experiments that inspired it to the three settings that control it. Lab: walk a graph by hand, then tune HNSW for recall against speed.
- Chapter 6: Billion Scale. The memory math that decides your hardware bill, scalar, binary, and product quantization, rescoring, and DiskANN. Lab: compress vectors four ways and measure what each costs in recall.
- Chapter 7: From Library to Database. What a database adds on top of an index, how to check its recall in production, and how to choose between a vector extension like pgvector and a purpose-built vector database. Lab: pgvector in Docker, with filtered search and a recall check.
- Chapter 8: Hybrid Search. BM25, sparse vectors, and rank fusion, and why vectors alone miss error codes and names. Lab: dense vs. keyword vs. fused search.
- Chapter 9: Rerankers. Bi-encoders, cross-encoders, and late interaction, and what a second pass costs. Lab: measure quality and latency before and after reranking.
- Chapter 10: GraphRAG on a Vector Store. Questions that need two hops, an entity graph stored next to your vectors, and what Microsoft GraphRAG adds. Lab: find the second paragraph a two-hop question needs, which vector search alone misses.
- Chapter 11: Capstone. A full retrieval stack with evaluation and a UI to inspect every stage.
What's next​
Do the Setup page first. Then Chapter 1 goes back to where this started: a search system that could only match exact words, and the idea that fixed it by turning text into points in space.