Chapter 1: Meaning as Geometry
Time: 25 minutes. Cost: $0 with Ollama, a fraction of a cent with OpenAI.
- Search that matches words fails when people use different words for the same thing, like "laptop" and "notebook."
- The fix took about 70 years to build: turn every text into a list of numbers, a point in space, so that texts with similar meaning land close together.
- The lab runs one question three ways, from word counting to a modern embedding model, and the right answer climbs from fourth place to first.
An employee types my laptop won't turn on into the company help desk. The best article in the
knowledge base begins "A notebook computer that will not start..." A search that matches words
finds nothing in common, because there is not anything: not one word overlaps.
This is the vocabulary mismatch problem: people describe the same thing with different words. Every idea in this chapter is an attempt to solve it. They all build on one older idea, the one this whole track rests on: a document can be treated as a point in space.
Documents as points​
Suppose you only cared about two words, "laptop" and "printer." Count how often each article uses them, and you can place every article on a sheet of graph paper: laptop count across, printer count up. Here are three tiny articles and a query:
| Text | "laptop" | "printer" |
|---|---|---|
| A: "My laptop will not charge. The laptop light stays off." | 2 | 0 |
| B: "The printer jams when the laptop sends a big file." | 1 | 1 |
| C: "The printer is out of toner. Replace the printer cartridge." | 0 | 2 |
| Query: "laptop not charging" | 1 | 0 |
Each arrow starts at the corner, where both counts are zero. The query's arrow points straight along the laptop direction. So does A's, which is longer because A says "laptop" twice. B's arrow points halfway between laptop and printer. C's points straight up, at printer. Measured by the angle between arrows, A is the closest match, B is a partial match, and C has nothing to do with the query. Comparing by that angle is the cosine similarity you met in Foundations.
That is the whole idea. A real collection has thousands of words instead of two, so each article is a point in a space with thousands of directions instead of a flat sheet. You cannot draw it, but the arithmetic is the same. A query is just a short document, so it becomes a point in the same space, and search means finding the documents that sit closest to it.
Gerard Salton and his SMART retrieval system turned word counting into this geometry in the 1960s. Vector databases still run on his idea.
Where this came from: Luhn, Salton, and the vector space model
In 1957, Hans Peter Luhn at IBM proposed indexing documents statistically, by counting their words. His intuition was simple: a paper that uses "transistor" thirty times is probably about transistors.
Salton's SMART system, started at Harvard and developed at Cornell from 1965, treated every distinct word in a collection as a direction in space, exactly like the two-word picture above. His group compared query and document vectors by the angle between them. The idea got its best-known name in a 1975 paper, "A Vector Space Model for Automatic Indexing" by Salton, Wong, and Yang.
Better weights: TF-IDF​
Raw counts have a flaw. On a help desk, nearly every ticket says "please" and "help." Only two say "BitLocker." If a new ticket mentions BitLocker, that one word tells you more than the rest of the sentence put together. A word that is everywhere is no clue at all.
So weight each word by how rare it is across the collection. That weight is called inverse document frequency (IDF). Multiply how often a word appears in this document (term frequency, TF) by how rare it is overall (IDF), and you get TF-IDF, a standard weighting in search systems for decades. Karen Spärck Jones published the rarity idea in 1972.
The formula, and who invented which half
The lab computes IDF as log(N / number of articles containing the word), where N is the number of
articles. A word in 2 of 16 articles gets log(8), about 2.1. A word in all 16 gets log(1), which is
exactly 0, so it adds nothing to any vector.
Spärck Jones introduced the IDF half in 1972. The combined TF-IDF weighting grew out of the SMART work that followed.
A TF-IDF vector is sparse: it has one number, or dimension, per vocabulary word, and almost all
of them are zero for any given document. In this chapter's lab, the vocabulary is 126 words, and the query
my laptop won't turn on has exactly one non-zero dimension: laptop. The lab also drops a short list
of stop words, very common words like "my" and "on" that IDF would weight near zero anyway.
"won't" and "turn" survive that list but appear in no article, so they get no dimension at all.
That is also where TF-IDF breaks. When every word is its own direction, "laptop" and "notebook" are as unrelated as "laptop" and "printer." The article about a notebook computer that will not start scores exactly zero.
Let the data find the concepts: LSA​
If two people keep turning up in the same group photos, you might guess they know each other, even if you never see them talk. In 1990, a team of five researchers published latent semantic analysis (LSA), which makes the same guess about words. Words that keep appearing in the same documents probably mean related things, so let the math find those groups.
LSA takes the big table of TF-IDF numbers, one row per document and one column per word, and squeezes it down to a handful of concept columns. If "laptop" and "notebook" show up in the same documents often enough, they end up in the same concepts. A query about laptops can then match a document about notebooks.
The catch is in "often enough." LSA only knows what your own documents tell it. On a big collection that works well. On the 16 help-desk articles in this chapter's lab, it is fragile: only one article uses both "laptop" and "notebook," and change the number of concept dimensions and the ranking shifts.
How LSA squeezes the table, and who published it
LSA was published by Deerwester, Dumais, Furnas, Landauer, and Harshman. The squeezing step uses a technique from linear algebra called singular value decomposition (SVD). SVD rewrites the word-by-document table as a set of directions ranked by how much of the table each one explains. LSA keeps the top hundred or so, or 4 in this chapter's lab, and throws the rest away. Words that are used the same way end up pointing in similar directions.
Learn meaning from everything ever written​
Fill in the blank: "I left my ___ charging overnight, and it still will not turn on."
Laptop, notebook, and phone all fit. Printer sounds odd. Banana does not fit at all. You know that because you have read thousands of sentences like it, and words that fit the same blanks tend to mean similar things. The linguist J.R. Firth put it memorably in 1957: "You shall know a word by the company it keeps." LSA used that idea on a few documents. Neural networks use it on a huge scale.
A neural network learns word meaning by playing that fill-in-the-blank game billions of times, on more text than a person could read in many lifetimes. To get good at the game, it has to give "laptop" and "notebook" similar numbers, because they fit the same blanks. In 2013, Google's word2vec made learning word vectors this way cheap enough to run on very large text collections.
Those vectors are dense, and the difference from TF-IDF is worth a picture. A TF-IDF vector is like a list of ingredients: one line for every possible ingredient, almost all of them empty. A dense vector is like a taste profile: a few hundred scores, every one of them filled in. Two dishes can taste alike with no ingredient in common, the same way "my laptop won't turn on" and "a notebook computer that will not start" mean the same thing with no word in common. The difference from a real taste profile is that nobody names the scores. There is no "sweet" column, and no column that means "laptop." Meaning is spread across all of them.
Word vectors still had to be combined to represent a whole sentence. Sentence-BERT, in 2019, trained
a model to produce one vector per sentence, so that sentences with similar meanings land close together
and can be compared with plain cosine similarity. Its paper made the cost argument plainly: finding
the most similar pair in 10,000 sentences took about 65 hours when a model had to read every pair together,
and about 5 seconds with Sentence-BERT. The embedding models you call today, including nomic-embed-text in
this lab, descend from that approach.
Where this came from: from Harris in 1954 to Sentence-BERT
The linguist Zellig Harris argued in 1954 that words with similar meanings occur in similar contexts. Firth's line came three years later. Together they are called the distributional hypothesis.
In the early 2000s, Yoshua Bengio and colleagues trained a neural network to predict the next word, and it learned a vector for every word as part of that job. Tomas Mikolov's team at Google published word2vec in 2013, followed by Stanford's GloVe in 2014.
By 2018, Google's BERT, a transformer (the neural network design behind modern language models), could judge whether two sentences meant the same thing, but only by reading both together, one pair at a time. That is what made comparing 10,000 sentences take 65 hours. Sentence-BERT (Reimers and Gurevych, 2019) retrained BERT to produce one vector per sentence instead.
The last push came in 2020, when a team at Facebook AI Research published retrieval-augmented generation (RAG). Think of an open-book exam. Instead of answering from memory, the language model first looks up the relevant pages, then answers with those pages in front of it. Vector search is how it finds the pages. RAG gave vector search the biggest job it has ever had.
The pipeline never changes: text in, vector out, compare by angle. Seventy years of progress went into the middle box, deciding what the vector should be.
Sparse is not dead​
It would be easy to read this chapter as "embeddings won, word counting lost." That is not what
happened. A sparse vector matches exact words, and sometimes exact words are what you need: an
error code like 0x80070005, a product name, a person's name. Think of a stock number on a part.
"Something like part 4471" is no use when you need exactly 4471. An embedding model may place an
error code near other error codes, which is exactly wrong when you need that one. Chapter 8 brings
sparse vectors back and combines them with dense ones. For now, hold on to the distinction:
| Sparse (TF-IDF) | Dense (embeddings) | |
|---|---|---|
| Like | A list of ingredients | A taste profile |
| Dimensions | One per vocabulary word, thousands to millions | Fixed, usually a few hundred to a few thousand |
| Non-zero values | A handful per document | All of them |
| What a dimension means | A specific word | Nothing you can name |
| Good at | Exact terms, codes, names | Synonyms, paraphrases, meaning |
| Blind to | Synonyms | Exact rare tokens it never learned |
Hands-on lab: one search, three generations​
You will run the help-desk query against 16 articles three ways: TF-IDF built by hand in plain Python, LSA built from that TF-IDF table, and a modern embedding model. All three rank articles by cosine similarity. Only the vectors change.
One detail in the code: nomic-embed-text was trained with a short label in front of every text,
search_query: for questions and search_document: for the articles being searched, and its model
card asks you to keep them. The lab adds them for you. Many embedding models have a convention like
this, and skipping it quietly costs you accuracy.
Full instructions: download the
Vector Databases labs ZIP, then open
labs/vector-databases/01-meaning-as-geometry and follow its README.
A real run, with Ollama and nomic-embed-text:
Query: "my laptop won't turn on"
TF-IDF vectors have 126 dimensions, one per word.
Query words the articles also use: ['laptop']
Non-zero dimensions in the query vector: 1 of 126
1. TF-IDF (matching words)
1. (0.24) Your notebook or laptop can be locked from the IT portal if it is lost or stolen.
2. (0.17) To connect a second monitor, use the USB-C port on the left side of the laptop.
3. (0.16) Laptop batteries last longer if you avoid running them all the way down to zero.
2. LSA (4 concept dimensions learned from these 16 articles)
1. (0.96) If your laptop does not power up, hold the power button for 30 seconds, then plug in the charger and try again.
2. (0.85) Laptop batteries last longer if you avoid running them all the way down to zero.
3. (0.69) To connect a second monitor, use the USB-C port on the left side of the laptop.
3. Embedding model (768 dimensions, learned from huge amounts of text)
1. (0.78) If your laptop does not power up, hold the power button for 30 seconds, then plug in the charger and try again.
2. (0.70) A notebook computer that will not start often has a drained battery. Leave it on the charger for an hour.
3. (0.61) To connect a second monitor, use the USB-C port on the left side of the laptop.
Spotlight: "A notebook computer that will not start often has a drained battery. Leave it on the charger for an hour."
It shares no words with the query. How each method scores it:
TF-IDF 0.00
LSA 0.37
Embedding 0.70
Read it top to bottom. TF-IDF found the word "laptop" and nothing else, so every laptop article looks about the same to it. The one that actually answers the question came fourth, because cosine divides by the article's length: that article is long, so its single "laptop" counts for less.
LSA put the right answer first and gave the notebook article some credit despite zero shared words. Be skeptical of that win, though. The lab uses 4 concept dimensions because that is where LSA does best on these 16 articles. At 5, the monitor article takes first place; at 6 or more, the battery article does, and the notebook article's score drops to about 0.1. On a collection this small, LSA's concepts are mostly noise. The README's first exercise has you try this.
The embedding model ranked both answers first and second, with no tuning, because it learned long before this lab that a notebook computer that will not start and a laptop that won't turn on are the same problem.
Checkpoint​
In the vector space model, what is a query, and how is it compared to documents?
A query is treated as a short document and turned into a vector in the same space as the documents. Search ranks documents by how close their vectors are to the query's, classically by the cosine of the angle between them.
Why did the "notebook computer that will not start" article score exactly 0.00 with TF-IDF?
In TF-IDF, every word is its own dimension. That article and the query share no words, so wherever one vector is non-zero the other is zero. Their dot product, and so their cosine, is exactly zero.
What is the difference between how LSA and an embedding model learn that "laptop" and "notebook" are related?
LSA learns it only from your own documents, from words that appear together in them, so it needs a large collection and is fragile on a small one. An embedding model learned it beforehand from a huge amount of general text, so it works even on 16 articles it has never seen.
Check Your Knowledge​
Click to start the quiz
What's next​
Every method in this chapter ranked articles by cosine similarity, and Foundations told you cosine is the safe default. Chapter 2 asks why. It compares cosine, dot product, and Euclidean distance, shows exactly when they disagree, and measures what happens to "closest" when you have 768 dimensions instead of 2.