Skip to main content

Chapter 1: Meaning as Geometry

Time: 25 minutes. Cost: $0 with Ollama, a fraction of a cent with OpenAI.

The short version
  • Search that matches words fails when people use different words for the same thing, like "laptop" and "notebook."
  • The fix took about 70 years to build: turn every text into a list of numbers, a point in space, so that texts with similar meaning land close together.
  • The lab runs one question three ways, from word counting to a modern embedding model, and the right answer climbs from fourth place to first.

An employee types my laptop won't turn on into the company help desk. The best article in the knowledge base begins "A notebook computer that will not start..." A search that matches words finds nothing in common, because there is not anything: not one word overlaps.

This is the vocabulary mismatch problem: people describe the same thing with different words. Every idea in this chapter is an attempt to solve it. They all build on one older idea, the one this whole track rests on: a document can be treated as a point in space.

Documents as points​

Suppose you only cared about two words, "laptop" and "printer." Count how often each article uses them, and you can place every article on a sheet of graph paper: laptop count across, printer count up. Here are three tiny articles and a query:

Text"laptop""printer"
A: "My laptop will not charge. The laptop light stays off."20
B: "The printer jams when the laptop sends a big file."11
C: "The printer is out of toner. Replace the printer cartridge."02
Query: "laptop not charging"10
Three articles and a query drawn as arrowsLaptop count runs across and printer count runs up. The query arrow points along the laptop axis to (1, 0). Article A points the same way, twice as far, to (2, 0). Article B points halfway between, to (1, 1). Article C points straight up the printer axis, to (0, 2).1212"laptop" count"printer" countABCQuery
The query, in amber, is drawn just above the axis so you can see it next to A.

Each arrow starts at the corner, where both counts are zero. The query's arrow points straight along the laptop direction. So does A's, which is longer because A says "laptop" twice. B's arrow points halfway between laptop and printer. C's points straight up, at printer. Measured by the angle between arrows, A is the closest match, B is a partial match, and C has nothing to do with the query. Comparing by that angle is the cosine similarity you met in Foundations.

That is the whole idea. A real collection has thousands of words instead of two, so each article is a point in a space with thousands of directions instead of a flat sheet. You cannot draw it, but the arithmetic is the same. A query is just a short document, so it becomes a point in the same space, and search means finding the documents that sit closest to it.

Gerard Salton and his SMART retrieval system turned word counting into this geometry in the 1960s. Vector databases still run on his idea.

Where this came from: Luhn, Salton, and the vector space model

In 1957, Hans Peter Luhn at IBM proposed indexing documents statistically, by counting their words. His intuition was simple: a paper that uses "transistor" thirty times is probably about transistors.

Salton's SMART system, started at Harvard and developed at Cornell from 1965, treated every distinct word in a collection as a direction in space, exactly like the two-word picture above. His group compared query and document vectors by the angle between them. The idea got its best-known name in a 1975 paper, "A Vector Space Model for Automatic Indexing" by Salton, Wong, and Yang.

Better weights: TF-IDF​

Raw counts have a flaw. On a help desk, nearly every ticket says "please" and "help." Only two say "BitLocker." If a new ticket mentions BitLocker, that one word tells you more than the rest of the sentence put together. A word that is everywhere is no clue at all.

So weight each word by how rare it is across the collection. That weight is called inverse document frequency (IDF). Multiply how often a word appears in this document (term frequency, TF) by how rare it is overall (IDF), and you get TF-IDF, a standard weighting in search systems for decades. Karen Spärck Jones published the rarity idea in 1972.

The formula, and who invented which half

The lab computes IDF as log(N / number of articles containing the word), where N is the number of articles. A word in 2 of 16 articles gets log(8), about 2.1. A word in all 16 gets log(1), which is exactly 0, so it adds nothing to any vector.

Spärck Jones introduced the IDF half in 1972. The combined TF-IDF weighting grew out of the SMART work that followed.

A TF-IDF vector is sparse: it has one number, or dimension, per vocabulary word, and almost all of them are zero for any given document. In this chapter's lab, the vocabulary is 126 words, and the query my laptop won't turn on has exactly one non-zero dimension: laptop. The lab also drops a short list of stop words, very common words like "my" and "on" that IDF would weight near zero anyway. "won't" and "turn" survive that list but appear in no article, so they get no dimension at all.

That is also where TF-IDF breaks. When every word is its own direction, "laptop" and "notebook" are as unrelated as "laptop" and "printer." The article about a notebook computer that will not start scores exactly zero.

Let the data find the concepts: LSA​

If two people keep turning up in the same group photos, you might guess they know each other, even if you never see them talk. In 1990, a team of five researchers published latent semantic analysis (LSA), which makes the same guess about words. Words that keep appearing in the same documents probably mean related things, so let the math find those groups.

LSA takes the big table of TF-IDF numbers, one row per document and one column per word, and squeezes it down to a handful of concept columns. If "laptop" and "notebook" show up in the same documents often enough, they end up in the same concepts. A query about laptops can then match a document about notebooks.

The catch is in "often enough." LSA only knows what your own documents tell it. On a big collection that works well. On the 16 help-desk articles in this chapter's lab, it is fragile: only one article uses both "laptop" and "notebook," and change the number of concept dimensions and the ranking shifts.

How LSA squeezes the table, and who published it

LSA was published by Deerwester, Dumais, Furnas, Landauer, and Harshman. The squeezing step uses a technique from linear algebra called singular value decomposition (SVD). SVD rewrites the word-by-document table as a set of directions ranked by how much of the table each one explains. LSA keeps the top hundred or so, or 4 in this chapter's lab, and throws the rest away. Words that are used the same way end up pointing in similar directions.

Learn meaning from everything ever written​

Fill in the blank: "I left my ___ charging overnight, and it still will not turn on."

Laptop, notebook, and phone all fit. Printer sounds odd. Banana does not fit at all. You know that because you have read thousands of sentences like it, and words that fit the same blanks tend to mean similar things. The linguist J.R. Firth put it memorably in 1957: "You shall know a word by the company it keeps." LSA used that idea on a few documents. Neural networks use it on a huge scale.

A neural network learns word meaning by playing that fill-in-the-blank game billions of times, on more text than a person could read in many lifetimes. To get good at the game, it has to give "laptop" and "notebook" similar numbers, because they fit the same blanks. In 2013, Google's word2vec made learning word vectors this way cheap enough to run on very large text collections.

Those vectors are dense, and the difference from TF-IDF is worth a picture. A TF-IDF vector is like a list of ingredients: one line for every possible ingredient, almost all of them empty. A dense vector is like a taste profile: a few hundred scores, every one of them filled in. Two dishes can taste alike with no ingredient in common, the same way "my laptop won't turn on" and "a notebook computer that will not start" mean the same thing with no word in common. The difference from a real taste profile is that nobody names the scores. There is no "sweet" column, and no column that means "laptop." Meaning is spread across all of them.

Word vectors still had to be combined to represent a whole sentence. Sentence-BERT, in 2019, trained a model to produce one vector per sentence, so that sentences with similar meanings land close together and can be compared with plain cosine similarity. Its paper made the cost argument plainly: finding the most similar pair in 10,000 sentences took about 65 hours when a model had to read every pair together, and about 5 seconds with Sentence-BERT. The embedding models you call today, including nomic-embed-text in this lab, descend from that approach.

Where this came from: from Harris in 1954 to Sentence-BERT

The linguist Zellig Harris argued in 1954 that words with similar meanings occur in similar contexts. Firth's line came three years later. Together they are called the distributional hypothesis.

In the early 2000s, Yoshua Bengio and colleagues trained a neural network to predict the next word, and it learned a vector for every word as part of that job. Tomas Mikolov's team at Google published word2vec in 2013, followed by Stanford's GloVe in 2014.

By 2018, Google's BERT, a transformer (the neural network design behind modern language models), could judge whether two sentences meant the same thing, but only by reading both together, one pair at a time. That is what made comparing 10,000 sentences take 65 hours. Sentence-BERT (Reimers and Gurevych, 2019) retrained BERT to produce one vector per sentence instead.

The last push came in 2020, when a team at Facebook AI Research published retrieval-augmented generation (RAG). Think of an open-book exam. Instead of answering from memory, the language model first looks up the relevant pages, then answers with those pages in front of it. Vector search is how it finds the pages. RAG gave vector search the biggest job it has ever had.

The pipeline never changes: text in, vector out, compare by angle. Seventy years of progress went into the middle box, deciding what the vector should be.

Sparse is not dead​

It would be easy to read this chapter as "embeddings won, word counting lost." That is not what happened. A sparse vector matches exact words, and sometimes exact words are what you need: an error code like 0x80070005, a product name, a person's name. Think of a stock number on a part. "Something like part 4471" is no use when you need exactly 4471. An embedding model may place an error code near other error codes, which is exactly wrong when you need that one. Chapter 8 brings sparse vectors back and combines them with dense ones. For now, hold on to the distinction:

Sparse (TF-IDF)Dense (embeddings)
LikeA list of ingredientsA taste profile
DimensionsOne per vocabulary word, thousands to millionsFixed, usually a few hundred to a few thousand
Non-zero valuesA handful per documentAll of them
What a dimension meansA specific wordNothing you can name
Good atExact terms, codes, namesSynonyms, paraphrases, meaning
Blind toSynonymsExact rare tokens it never learned

Hands-on lab: one search, three generations​

You will run the help-desk query against 16 articles three ways: TF-IDF built by hand in plain Python, LSA built from that TF-IDF table, and a modern embedding model. All three rank articles by cosine similarity. Only the vectors change.

One detail in the code: nomic-embed-text was trained with a short label in front of every text, search_query: for questions and search_document: for the articles being searched, and its model card asks you to keep them. The lab adds them for you. Many embedding models have a convention like this, and skipping it quietly costs you accuracy.

Full instructions: download the Vector Databases labs ZIP, then open labs/vector-databases/01-meaning-as-geometry and follow its README.

A real run, with Ollama and nomic-embed-text:

Query: "my laptop won't turn on"

TF-IDF vectors have 126 dimensions, one per word.
Query words the articles also use: ['laptop']
Non-zero dimensions in the query vector: 1 of 126

1. TF-IDF (matching words)
1. (0.24) Your notebook or laptop can be locked from the IT portal if it is lost or stolen.
2. (0.17) To connect a second monitor, use the USB-C port on the left side of the laptop.
3. (0.16) Laptop batteries last longer if you avoid running them all the way down to zero.

2. LSA (4 concept dimensions learned from these 16 articles)
1. (0.96) If your laptop does not power up, hold the power button for 30 seconds, then plug in the charger and try again.
2. (0.85) Laptop batteries last longer if you avoid running them all the way down to zero.
3. (0.69) To connect a second monitor, use the USB-C port on the left side of the laptop.

3. Embedding model (768 dimensions, learned from huge amounts of text)
1. (0.78) If your laptop does not power up, hold the power button for 30 seconds, then plug in the charger and try again.
2. (0.70) A notebook computer that will not start often has a drained battery. Leave it on the charger for an hour.
3. (0.61) To connect a second monitor, use the USB-C port on the left side of the laptop.

Spotlight: "A notebook computer that will not start often has a drained battery. Leave it on the charger for an hour."
It shares no words with the query. How each method scores it:
TF-IDF 0.00
LSA 0.37
Embedding 0.70

Read it top to bottom. TF-IDF found the word "laptop" and nothing else, so every laptop article looks about the same to it. The one that actually answers the question came fourth, because cosine divides by the article's length: that article is long, so its single "laptop" counts for less.

LSA put the right answer first and gave the notebook article some credit despite zero shared words. Be skeptical of that win, though. The lab uses 4 concept dimensions because that is where LSA does best on these 16 articles. At 5, the monitor article takes first place; at 6 or more, the battery article does, and the notebook article's score drops to about 0.1. On a collection this small, LSA's concepts are mostly noise. The README's first exercise has you try this.

The embedding model ranked both answers first and second, with no tuning, because it learned long before this lab that a notebook computer that will not start and a laptop that won't turn on are the same problem.

Checkpoint​

In the vector space model, what is a query, and how is it compared to documents?

A query is treated as a short document and turned into a vector in the same space as the documents. Search ranks documents by how close their vectors are to the query's, classically by the cosine of the angle between them.

Why did the "notebook computer that will not start" article score exactly 0.00 with TF-IDF?

In TF-IDF, every word is its own dimension. That article and the query share no words, so wherever one vector is non-zero the other is zero. Their dot product, and so their cosine, is exactly zero.

What is the difference between how LSA and an embedding model learn that "laptop" and "notebook" are related?

LSA learns it only from your own documents, from words that appear together in them, so it needs a large collection and is fragile on a small one. An embedding model learned it beforehand from a huge amount of general text, so it works even on 16 articles it has never seen.

Check Your Knowledge​

Click to start the quiz
1. Suppose every one of the 16 help-desk articles contained the word "employee". What weight would the lab's TF-IDF give "employee" in any vector?
2. An employee pastes the exact error code 0x80070005 into the search box, and one article mentions that code. Which representation is most likely to rank that article first?
3. Which of the 768 dimensions in a nomic-embed-text vector represents the word "laptop"?
4. The lab ranks articles with the same cosine similarity function for TF-IDF, LSA, and the embedding model. Why build it that way?

What's next​

Every method in this chapter ranked articles by cosine similarity, and Foundations told you cosine is the safe default. Chapter 2 asks why. It compares cosine, dot product, and Euclidean distance, shows exactly when they disagree, and measures what happens to "closest" when you have 768 dimensions instead of 2.