Chapter 2: The Math of Similarity
Time: 30 minutes. Cost: $0 with Ollama, a fraction of a cent with OpenAI.
- There are three common ways to measure "closest": by direction only (cosine), by direction and size together (dot product), or by straight-line distance (Euclidean). They can disagree.
- For text, they usually agree, because many embedding models, including the two in this track, return vectors of length 1. Check that once, then use cosine or dot product.
- When size carries meaning, like a box's measurements or an item's popularity, the choice decides the answer. A table in this chapter maps common jobs to the right measure.
Here are three points on a flat sheet of paper: a query q = (2, 1), and two candidates,
a = (4, 2) and b = (2, 1.5). Which candidate is closest to q?
It depends on how you measure. In this chapter's lab, the dot product picks a. Cosine similarity
picks a. Euclidean distance picks b. Same three points, two different answers, and both are
correct by their own definition.
When you create a vector index, the database asks you to choose one of those measures. This chapter gives you the math to choose on purpose, in plain arithmetic, no calculus. Then it covers two things that are easy to get wrong: what happens to "closest" in 768 dimensions, and why a similarity score of 0.5 does not mean "half similar."
A vector is an arrow​
Picture each vector as an arrow from the origin (0, 0) to its point. Every arrow has two
properties: a direction and a length.
A recipe makes the difference concrete. Read q = (2, 1) as 2 cups of flour and 1 cup of sugar.
Then a = (4, 2) is the same recipe, doubled: same direction, twice the length. b = (2, 1.5) is
a slightly sweeter recipe at about the same size: a slightly different direction, about the same
length. Direction is what the recipe is. Length is how much of it you made.
Length is Pythagoras. For q = (2, 1), it is the square root of 2² + 1², about 2.24. In 768
dimensions it is the same rule with 768 squares to add instead of two:
length(v) = sqrt(v1² + v2² + ... + vn²)
The three measures disagree because they weigh direction and length differently.
Three measures​
Dot product. Multiply matching numbers and add them up:
dot(q, a) = 2×4 + 1×2 = 10
dot(q, b) = 2×2 + 1×1.5 = 5.5
The dot product grows when the arrows point the same way, and it also grows when either arrow gets
longer. Higher means closer. That is why it picks a: same recipe, and a bigger batch. (For the
curious: it equals length(q) × length(a) × cos(angle between them).)
Cosine similarity. Divide the dot product by both lengths, which cancels length out and leaves only the angle:
cosine(q, a) = dot(q, a) / (length(q) × length(a))
It ranges from -1 (opposite directions) through 0 (perpendicular) to 1 (same direction). a scores
exactly 1.00 because it points exactly where q points. In recipe terms, cosine asks "is this the
same recipe?" and does not care about the batch size.
Euclidean (L2) distance. Lay a ruler between the two points:
L2(q, b) = sqrt((2-2)² + (1-1.5)²) = 0.5
L2(q, a) = sqrt((2-4)² + (1-2)²) = 2.24
Lower means closer. b sits right next to q, while a overshoots by a whole q-length. In
recipe terms, L2 asks "how much would I have to add or remove to turn one bowl into the other?" Half
a cup of sugar beats doubling everything, so L2 picks b.
Normalize, and the argument disappears​
Normalizing a vector means dividing it by its own length, so that its length becomes exactly 1.
The direction does not change. It is like scaling every recipe to make exactly one cup. Normalize
q, a, and b and run the three measures again: all three pick a.
That is not luck. Once every arrow has length 1, length cannot make a difference, so only the angle is left, and all three measures produce exactly the same ranking. The lab checks this on real embeddings: across all 16 articles, the three rankings match exactly.
The two identities that make it work
On unit-length vectors:
dot(u, v) = cosine(u, v) because both lengths are 1
L2(u, v)² = 2 - 2 × cosine(u, v) expand (u - v)·(u - v) = 1 - 2·u·v + 1
Higher cosine always means lower L2 distance, and dot product is cosine. The lab checks the second
identity on every article: the largest gap between L2² and 2 - 2 × cosine is 0.0000005, which is
just floating-point rounding.
This is why many production systems normalize once, when the vector is stored, and then use the dot product. It is one multiply-and-add per dimension, with no division by lengths at query time.
Check your vectors, do not assume​
Whether your vectors are already normalized depends on the provider, sometimes on which endpoint
you call. OpenAI documents that its embeddings are normalized to length 1. Ollama documents the same
for its current /api/embed endpoint, and our runs agree: nomic-embed-text vectors came back at
length 1.0000. Its older /api/embeddings endpoint, which Ollama's docs say is
superseded, returned vectors from the same model at lengths of about 20 to 21. (The lab README
shows how to reproduce that.)
Same model, same text, two endpoints, and only one of them gives you vectors where dot product equals
cosine. Checking is one line of numpy, np.linalg.norm(vector), and the lab prints it for you.
Which measure for which job​
For text embeddings from a standard model, the answer is short: the model put the meaning in the direction, so use cosine, or dot product on normalized vectors. Length matters when you put it there on purpose, or when your numbers are not embeddings at all. Here is how that plays out:
| Job | What the vectors hold | Use | Why it fits |
|---|---|---|---|
| Search support tickets or help articles | Text embeddings | Cosine, or dot product if lengths are 1 | A two-line ticket and a two-page ticket about the same problem point the same way. Length says nothing about the topic. |
| Search a product catalog by description | Text embeddings | Cosine, or dot product if lengths are 1 | Same reason. "Waterproof hiking boots" should match boots described in a sentence or a paragraph. |
| Find duplicate tickets | Text embeddings | Cosine, with a cutoff you set from labeled pairs | Same direction means same content. The cutoff decides how alike counts as a duplicate. The last section of this chapter explains why it has to come from your own data. |
| Find stock items by size, like boxes or parts | Measurements you took, in the same units | Euclidean | The size is the point. Cosine thinks a shoebox and a crate with the same proportions are identical. The lab shows this. |
| Recommend items, when popularity should count | Item vectors from a recommendation model | Dot product, which the literature calls maximum inner product search | Google's recommendation systems course notes that popular items tend to get longer vectors. Dot product keeps that signal. Too much of it, and popular items crowd out everything else. |
| Search images by text, or anything from one specific model | Whatever that model produces | The measure the model was trained with | CLIP, a widely used image and text model, was trained to compare by cosine. The model card tells you. |
The box row is worth seeing with real numbers. An order needs a box about 10 × 20 × 30 cm, and the warehouse stores every box as three measurements:
PART 5: an order needs a box about 10 x 20 x 30 cm
dot cosine L2
small box 12 x 18 x 30 cm 1380 0.997 2.8
tall box 10 x 20 x 60 cm 2300 0.960 30.0
crate 100 x 200 x 300 cm 14000 1.000 336.7
Best by dot: crate 100 x 200 x 300 cm
Best by cosine: crate 100 x 200 x 300 cm
Best by L2: small box 12 x 18 x 30 cm
The crate is the order's shape scaled up ten times, the doubled recipe again. Cosine gives it a perfect 1.000 and would ship a shoebox-sized order in a crate you could sleep in. Dot product likes it even more, because it is big. Only L2 picks the small box, because only L2 cares how far apart the actual sizes are.
The flowchart sums up the decision:
Check your own case​
Before you create an index, walk through these five questions:
- Where do the vectors come from? If a model made them, read its model card and use the measure it names. Most text embedding models expect cosine.
- Are they length 1? Print
np.linalg.normfor a few of them. If they are, all three measures rank the same way, so pick dot product, the cheapest. - If not, does a bigger number mean something? If it is a strength you want to reward, like popularity, use dot product. If it is an amount you want to match, like a size or a price, use L2. If only the mix matters, use cosine.
- For L2 on numbers you measured, are they on the same scale? If one column is price in dollars and another is weight in grams, a 1-gram difference counts as much as a 1-dollar one, and the column with the biggest numbers decides everything. Rescale each column first, for example to a 0-to-1 range.
- Does it work on your data? Write down 20 real searches and the result you would want for each. Then check which measure puts that result near the top. Chapter 3 calls a list like that ground truth, and it is the most useful thing you will build in this track.
Decide before you build the index. In many databases, pgvector included, an index is built for one measure, and switching means building it again.
The curse of dimensionality​
Pick two strangers and compare them on one thing, say height. Some pairs are almost the same height, and others are a foot apart. Now compare them on 768 unrelated things: height, shoe size, commute time, number of cousins, and so on. Every pair matches on some and differs on others, and the differences average out. Every pair of strangers ends up about equally different.
That is the strange form the curse of dimensionality takes for nearest-neighbor search. As the number of dimensions grows, the distance to the nearest point and the distance to the farthest point become nearly the same.
The lab measures it. Scatter 1,000 random points in a cube, pick a random query, and compare the nearest and farthest distances. Contrast is how much farther the farthest point is, relative to the nearest:
dimensions nearest farthest contrast
2 0.02 0.90 42.52
10 0.55 2.02 2.63
100 3.27 4.69 0.43
768 10.55 12.14 0.15
In 2 dimensions, the farthest point is 42 times farther than the nearest. In 768 dimensions, it is 15% farther. Everything is roughly equally far from everything else. If real embeddings behaved like this, "nearest neighbor" would be close to meaningless.
They do not, and the lab shows that too. It takes 16 points in 768 dimensions, lets each point take a turn as the query against the other 15, and averages the contrast. Sixteen random unit vectors average 0.06. The 16 real help-desk articles average 0.29, about five times higher. Real text is not a set of strangers with random traits. Meaning occupies a much smaller, structured region of that space, so near things are still meaningfully nearer.
The curse still matters, in two ways. It is why the tree indexes in Chapter 3 fail, and why every fast index in this track is approximate. And it is a warning about synthetic benchmarks: random vectors are much harder to search than real ones, so a benchmark on random data tells you little about your own.
Where the name came from
Richard Bellman coined the phrase "curse of dimensionality" in 1957, in a book on dynamic programming, to describe how problems explode as you add variables. The nearest-neighbor version came later: Beyer, Goldstein, Ramakrishnan, and Shaft showed in 1999 that for a broad class of data distributions, as the number of dimensions grows, the nearest and farthest distances become nearly the same.
What similarity scores really look like​
Foundations said cosine similarity runs from -1 to 1, with 0 meaning unrelated. That is true of the math. It is not what you see from a real model. The lab scores every pair of different help-desk articles against each other:
0.4 to 0.5 | ###
0.5 to 0.6 | ############################################
0.6 to 0.7 | #################################################################
0.7 to 0.8 | ########
120 pairs. Lowest 0.48, median 0.61, highest 0.79
Query vs. its best article: 0.78
(That is an excerpt: the lab also prints the empty bins and the best article's text.)
Articles about printers, passwords, Wi-Fi, and laptops, all on different topics, still score at least 0.48 against each other, and nothing comes anywhere near 0. The right answer to the query scores 0.78, slightly below the most similar pair of articles in the whole collection.
This model's vectors sit in a narrow cone instead of spreading across all directions: even
unrelated articles point in broadly similar directions, so the whole score range is squeezed into
the upper half. How narrow depends on the model. We ran the same 16 articles through a second
local model, mxbai-embed-large, using the query prompt its model card recommends (the lab README
shows how to try it):
| Model | Lowest pair | Median pair | Highest pair | Query vs. best article |
|---|---|---|---|---|
nomic-embed-text | 0.48 | 0.61 | 0.79 | 0.78 |
mxbai-embed-large | 0.26 | 0.44 | 0.83 | 0.81 |
Both models picked the same best article. The numbers around it are different.
Think of two teachers who grade differently. A 78 from a tough grader can be a better paper than an 81 from an easy one. You can rank the papers in one teacher's pile, but you cannot set one pass mark for both.
That has a practical consequence: a fixed similarity threshold does not transfer. A rule like
"only use results above 0.8" accepts the right answer with mxbai-embed-large, at 0.81, and rejects
it with nomic-embed-text, at 0.78. Scores are for ranking within one model on one collection. If
you need a cutoff, for example to answer "I do not know" when nothing relevant exists, set it from
labeled examples on your own data, with your own model, and re-check it whenever either changes.
Hands-on lab: measure all of it​
The lab works through five parts: the 2D disagreement, normalization on the 2D points and on real
embeddings, the curse of dimensionality, the score histogram, and the box order. Every formula from this chapter
appears in the code as a short function you can read: dot, length, cosine, and l2.
Full instructions: download the
Vector Databases labs ZIP, then open
labs/vector-databases/02-similarity-math and follow its README.
A real run, with Ollama:
PART 1: which point is closest to q?
q = [2. 1.], a = [4. 2.], b = [2. 1.5]
a b winner
dot 10.00 5.50 a (higher wins)
cosine 1.00 0.98 a (higher wins)
L2 2.24 0.50 b (lower wins)
PART 2: the same points, scaled to length 1
a b winner
dot 1.00 0.98 a
cosine 1.00 0.98 a
L2 0.00 0.18 a
Real embeddings: 16 articles, 768 dimensions each
Shortest vector length 1.0000, longest 1.0000
Top 5 by dot: [0, 1, 8, 3, 9]
Top 5 by cosine: [0, 1, 8, 3, 9]
Top 5 by L2: [0, 1, 8, 3, 9]
Largest gap between L2^2 and 2 - 2*cosine: 0.0000005
Parts 3, 4, and 5 print the tables and histogram shown earlier. The random points use a fixed seed and Part 5 is plain arithmetic, so those numbers should match yours exactly. Part 4 depends on the embedding model.
Checkpoint​
In Part 1, why does cosine similarity pick a while Euclidean distance picks b?
a points in exactly the same direction as q but is twice as long. Cosine ignores length and only
measures the angle, so a scores a perfect 1.00. Euclidean distance measures how far apart the two
points are, and a overshoots q by a whole q-length (2.24 away), while b is only 0.5 away.
Why do all three measures give the same ranking once vectors are normalized?
On unit-length vectors, the dot product equals the cosine, and L2² = 2 - 2 × cosine. A higher cosine always means a lower L2 distance, so all three order every candidate the same way.
The lab's 16 real articles average a contrast of 0.29 in 768 dimensions, while 16 random unit vectors average 0.06. What does that tell you?
Real embeddings are not spread randomly through all 768 dimensions. Meaning occupies a smaller, structured region, so the nearest article is still meaningfully closer than the farthest. Random vectors suffer the full curse of dimensionality. That is also why benchmarks on random vectors are poor predictors for real data.
A real estate site wants to show homes similar to one a buyer liked, using bedrooms, bathrooms, floor area in square meters, and price in dollars. Which measure, and what must it do first?
L2 distance, because the amounts are the point: a 2-bedroom flat and a 20-bedroom mansion with the same proportions are not similar homes, though cosine would call them close. First it must rescale each column. Otherwise price, in hundreds of thousands, swamps bedrooms, which run from 1 to 6, and the search becomes "homes at a similar price."
Check Your Knowledge​
Click to start the quiz
What's next​
Every search so far has compared the query against every single article, which is fine for 16. In Chapter 3, you will find out how far that brute-force approach goes, why the obvious speedup, a tree index, fails in high dimensions for exactly the reason you just measured, and where the nearest-neighbor idea came from in the first place.