Skip to main content

Chapter 2: The Math of Similarity

Time: 30 minutes. Cost: $0 with Ollama, a fraction of a cent with OpenAI.

The short version
  • There are three common ways to measure "closest": by direction only (cosine), by direction and size together (dot product), or by straight-line distance (Euclidean). They can disagree.
  • For text, they usually agree, because many embedding models, including the two in this track, return vectors of length 1. Check that once, then use cosine or dot product.
  • When size carries meaning, like a box's measurements or an item's popularity, the choice decides the answer. A table in this chapter maps common jobs to the right measure.

Here are three points on a flat sheet of paper: a query q = (2, 1), and two candidates, a = (4, 2) and b = (2, 1.5). Which candidate is closest to q?

It depends on how you measure. In this chapter's lab, the dot product picks a. Cosine similarity picks a. Euclidean distance picks b. Same three points, two different answers, and both are correct by their own definition.

When you create a vector index, the database asks you to choose one of those measures. This chapter gives you the math to choose on purpose, in plain arithmetic, no calculus. Then it covers two things that are easy to get wrong: what happens to "closest" in 768 dimensions, and why a similarity score of 0.5 does not mean "half similar."

A vector is an arrow​

Picture each vector as an arrow from the origin (0, 0) to its point. Every arrow has two properties: a direction and a length.

A recipe makes the difference concrete. Read q = (2, 1) as 2 cups of flour and 1 cup of sugar. Then a = (4, 2) is the same recipe, doubled: same direction, twice the length. b = (2, 1.5) is a slightly sweeter recipe at about the same size: a slightly different direction, about the same length. Direction is what the recipe is. Length is how much of it you made.

Length is Pythagoras. For q = (2, 1), it is the square root of 2² + 1², about 2.24. In 768 dimensions it is the same rule with 768 squares to add instead of two:

length(v) = sqrt(v1² + v2² + ... + vn²)

The three measures disagree because they weigh direction and length differently.

Three measures​

Dot product. Multiply matching numbers and add them up:

dot(q, a) = 2×4 + 1×2 = 10
dot(q, b) = 2×2 + 1×1.5 = 5.5

The dot product grows when the arrows point the same way, and it also grows when either arrow gets longer. Higher means closer. That is why it picks a: same recipe, and a bigger batch. (For the curious: it equals length(q) × length(a) × cos(angle between them).)

Cosine similarity. Divide the dot product by both lengths, which cancels length out and leaves only the angle:

cosine(q, a) = dot(q, a) / (length(q) × length(a))

It ranges from -1 (opposite directions) through 0 (perpendicular) to 1 (same direction). a scores exactly 1.00 because it points exactly where q points. In recipe terms, cosine asks "is this the same recipe?" and does not care about the batch size.

Euclidean (L2) distance. Lay a ruler between the two points:

L2(q, b) = sqrt((2-2)² + (1-1.5)²) = 0.5
L2(q, a) = sqrt((2-4)² + (1-2)²) = 2.24

Lower means closer. b sits right next to q, while a overshoots by a whole q-length. In recipe terms, L2 asks "how much would I have to add or remove to turn one bowl into the other?" Half a cup of sugar beats doubling everything, so L2 picks b.

Normalize, and the argument disappears​

Normalizing a vector means dividing it by its own length, so that its length becomes exactly 1. The direction does not change. It is like scaling every recipe to make exactly one cup. Normalize q, a, and b and run the three measures again: all three pick a.

That is not luck. Once every arrow has length 1, length cannot make a difference, so only the angle is left, and all three measures produce exactly the same ranking. The lab checks this on real embeddings: across all 16 articles, the three rankings match exactly.

The two identities that make it work

On unit-length vectors:

dot(u, v) = cosine(u, v) because both lengths are 1
L2(u, v)² = 2 - 2 × cosine(u, v) expand (u - v)·(u - v) = 1 - 2·u·v + 1

Higher cosine always means lower L2 distance, and dot product is cosine. The lab checks the second identity on every article: the largest gap between L2² and 2 - 2 × cosine is 0.0000005, which is just floating-point rounding.

This is why many production systems normalize once, when the vector is stored, and then use the dot product. It is one multiply-and-add per dimension, with no division by lengths at query time.

Check your vectors, do not assume​

Whether your vectors are already normalized depends on the provider, sometimes on which endpoint you call. OpenAI documents that its embeddings are normalized to length 1. Ollama documents the same for its current /api/embed endpoint, and our runs agree: nomic-embed-text vectors came back at length 1.0000. Its older /api/embeddings endpoint, which Ollama's docs say is superseded, returned vectors from the same model at lengths of about 20 to 21. (The lab README shows how to reproduce that.)

Same model, same text, two endpoints, and only one of them gives you vectors where dot product equals cosine. Checking is one line of numpy, np.linalg.norm(vector), and the lab prints it for you.

Which measure for which job​

For text embeddings from a standard model, the answer is short: the model put the meaning in the direction, so use cosine, or dot product on normalized vectors. Length matters when you put it there on purpose, or when your numbers are not embeddings at all. Here is how that plays out:

JobWhat the vectors holdUseWhy it fits
Search support tickets or help articlesText embeddingsCosine, or dot product if lengths are 1A two-line ticket and a two-page ticket about the same problem point the same way. Length says nothing about the topic.
Search a product catalog by descriptionText embeddingsCosine, or dot product if lengths are 1Same reason. "Waterproof hiking boots" should match boots described in a sentence or a paragraph.
Find duplicate ticketsText embeddingsCosine, with a cutoff you set from labeled pairsSame direction means same content. The cutoff decides how alike counts as a duplicate. The last section of this chapter explains why it has to come from your own data.
Find stock items by size, like boxes or partsMeasurements you took, in the same unitsEuclideanThe size is the point. Cosine thinks a shoebox and a crate with the same proportions are identical. The lab shows this.
Recommend items, when popularity should countItem vectors from a recommendation modelDot product, which the literature calls maximum inner product searchGoogle's recommendation systems course notes that popular items tend to get longer vectors. Dot product keeps that signal. Too much of it, and popular items crowd out everything else.
Search images by text, or anything from one specific modelWhatever that model producesThe measure the model was trained withCLIP, a widely used image and text model, was trained to compare by cosine. The model card tells you.

The box row is worth seeing with real numbers. An order needs a box about 10 × 20 × 30 cm, and the warehouse stores every box as three measurements:

PART 5: an order needs a box about 10 x 20 x 30 cm
dot cosine L2
small box 12 x 18 x 30 cm 1380 0.997 2.8
tall box 10 x 20 x 60 cm 2300 0.960 30.0
crate 100 x 200 x 300 cm 14000 1.000 336.7
Best by dot: crate 100 x 200 x 300 cm
Best by cosine: crate 100 x 200 x 300 cm
Best by L2: small box 12 x 18 x 30 cm

The crate is the order's shape scaled up ten times, the doubled recipe again. Cosine gives it a perfect 1.000 and would ship a shoebox-sized order in a crate you could sleep in. Dot product likes it even more, because it is big. Only L2 picks the small box, because only L2 cares how far apart the actual sizes are.

The flowchart sums up the decision:

Check your own case​

Before you create an index, walk through these five questions:

  1. Where do the vectors come from? If a model made them, read its model card and use the measure it names. Most text embedding models expect cosine.
  2. Are they length 1? Print np.linalg.norm for a few of them. If they are, all three measures rank the same way, so pick dot product, the cheapest.
  3. If not, does a bigger number mean something? If it is a strength you want to reward, like popularity, use dot product. If it is an amount you want to match, like a size or a price, use L2. If only the mix matters, use cosine.
  4. For L2 on numbers you measured, are they on the same scale? If one column is price in dollars and another is weight in grams, a 1-gram difference counts as much as a 1-dollar one, and the column with the biggest numbers decides everything. Rescale each column first, for example to a 0-to-1 range.
  5. Does it work on your data? Write down 20 real searches and the result you would want for each. Then check which measure puts that result near the top. Chapter 3 calls a list like that ground truth, and it is the most useful thing you will build in this track.

Decide before you build the index. In many databases, pgvector included, an index is built for one measure, and switching means building it again.

The curse of dimensionality​

Pick two strangers and compare them on one thing, say height. Some pairs are almost the same height, and others are a foot apart. Now compare them on 768 unrelated things: height, shoe size, commute time, number of cousins, and so on. Every pair matches on some and differs on others, and the differences average out. Every pair of strangers ends up about equally different.

That is the strange form the curse of dimensionality takes for nearest-neighbor search. As the number of dimensions grows, the distance to the nearest point and the distance to the farthest point become nearly the same.

The lab measures it. Scatter 1,000 random points in a cube, pick a random query, and compare the nearest and farthest distances. Contrast is how much farther the farthest point is, relative to the nearest:

dimensions nearest farthest contrast
2 0.02 0.90 42.52
10 0.55 2.02 2.63
100 3.27 4.69 0.43
768 10.55 12.14 0.15

In 2 dimensions, the farthest point is 42 times farther than the nearest. In 768 dimensions, it is 15% farther. Everything is roughly equally far from everything else. If real embeddings behaved like this, "nearest neighbor" would be close to meaningless.

They do not, and the lab shows that too. It takes 16 points in 768 dimensions, lets each point take a turn as the query against the other 15, and averages the contrast. Sixteen random unit vectors average 0.06. The 16 real help-desk articles average 0.29, about five times higher. Real text is not a set of strangers with random traits. Meaning occupies a much smaller, structured region of that space, so near things are still meaningfully nearer.

The curse still matters, in two ways. It is why the tree indexes in Chapter 3 fail, and why every fast index in this track is approximate. And it is a warning about synthetic benchmarks: random vectors are much harder to search than real ones, so a benchmark on random data tells you little about your own.

Where the name came from

Richard Bellman coined the phrase "curse of dimensionality" in 1957, in a book on dynamic programming, to describe how problems explode as you add variables. The nearest-neighbor version came later: Beyer, Goldstein, Ramakrishnan, and Shaft showed in 1999 that for a broad class of data distributions, as the number of dimensions grows, the nearest and farthest distances become nearly the same.

What similarity scores really look like​

Foundations said cosine similarity runs from -1 to 1, with 0 meaning unrelated. That is true of the math. It is not what you see from a real model. The lab scores every pair of different help-desk articles against each other:

0.4 to 0.5 | ###
0.5 to 0.6 | ############################################
0.6 to 0.7 | #################################################################
0.7 to 0.8 | ########

120 pairs. Lowest 0.48, median 0.61, highest 0.79
Query vs. its best article: 0.78

(That is an excerpt: the lab also prints the empty bins and the best article's text.)

Articles about printers, passwords, Wi-Fi, and laptops, all on different topics, still score at least 0.48 against each other, and nothing comes anywhere near 0. The right answer to the query scores 0.78, slightly below the most similar pair of articles in the whole collection.

This model's vectors sit in a narrow cone instead of spreading across all directions: even unrelated articles point in broadly similar directions, so the whole score range is squeezed into the upper half. How narrow depends on the model. We ran the same 16 articles through a second local model, mxbai-embed-large, using the query prompt its model card recommends (the lab README shows how to try it):

ModelLowest pairMedian pairHighest pairQuery vs. best article
nomic-embed-text0.480.610.790.78
mxbai-embed-large0.260.440.830.81

Both models picked the same best article. The numbers around it are different.

Think of two teachers who grade differently. A 78 from a tough grader can be a better paper than an 81 from an easy one. You can rank the papers in one teacher's pile, but you cannot set one pass mark for both.

That has a practical consequence: a fixed similarity threshold does not transfer. A rule like "only use results above 0.8" accepts the right answer with mxbai-embed-large, at 0.81, and rejects it with nomic-embed-text, at 0.78. Scores are for ranking within one model on one collection. If you need a cutoff, for example to answer "I do not know" when nothing relevant exists, set it from labeled examples on your own data, with your own model, and re-check it whenever either changes.

Hands-on lab: measure all of it​

The lab works through five parts: the 2D disagreement, normalization on the 2D points and on real embeddings, the curse of dimensionality, the score histogram, and the box order. Every formula from this chapter appears in the code as a short function you can read: dot, length, cosine, and l2.

Full instructions: download the Vector Databases labs ZIP, then open labs/vector-databases/02-similarity-math and follow its README.

A real run, with Ollama:

PART 1: which point is closest to q?
q = [2. 1.], a = [4. 2.], b = [2. 1.5]
a b winner
dot 10.00 5.50 a (higher wins)
cosine 1.00 0.98 a (higher wins)
L2 2.24 0.50 b (lower wins)

PART 2: the same points, scaled to length 1
a b winner
dot 1.00 0.98 a
cosine 1.00 0.98 a
L2 0.00 0.18 a

Real embeddings: 16 articles, 768 dimensions each
Shortest vector length 1.0000, longest 1.0000
Top 5 by dot: [0, 1, 8, 3, 9]
Top 5 by cosine: [0, 1, 8, 3, 9]
Top 5 by L2: [0, 1, 8, 3, 9]
Largest gap between L2^2 and 2 - 2*cosine: 0.0000005

Parts 3, 4, and 5 print the tables and histogram shown earlier. The random points use a fixed seed and Part 5 is plain arithmetic, so those numbers should match yours exactly. Part 4 depends on the embedding model.

Checkpoint​

In Part 1, why does cosine similarity pick a while Euclidean distance picks b?

a points in exactly the same direction as q but is twice as long. Cosine ignores length and only measures the angle, so a scores a perfect 1.00. Euclidean distance measures how far apart the two points are, and a overshoots q by a whole q-length (2.24 away), while b is only 0.5 away.

Why do all three measures give the same ranking once vectors are normalized?

On unit-length vectors, the dot product equals the cosine, and L2² = 2 - 2 × cosine. A higher cosine always means a lower L2 distance, so all three order every candidate the same way.

The lab's 16 real articles average a contrast of 0.29 in 768 dimensions, while 16 random unit vectors average 0.06. What does that tell you?

Real embeddings are not spread randomly through all 768 dimensions. Meaning occupies a smaller, structured region, so the nearest article is still meaningfully closer than the farthest. Random vectors suffer the full curse of dimensionality. That is also why benchmarks on random vectors are poor predictors for real data.

A real estate site wants to show homes similar to one a buyer liked, using bedrooms, bathrooms, floor area in square meters, and price in dollars. Which measure, and what must it do first?

L2 distance, because the amounts are the point: a 2-bedroom flat and a 20-bedroom mansion with the same proportions are not similar homes, though cosine would call them close. First it must rescale each column. Otherwise price, in hundreds of thousands, swamps bedrooms, which run from 1 to 6, and the search becomes "homes at a similar price."

Check Your Knowledge​

Click to start the quiz
1. A recommendation system deliberately trains item vectors so that more popular items get longer vectors, and it wants popularity to count when ranking. Which measure should it search with?
2. A teammate calls Ollama's older /api/embeddings endpoint, gets vectors with lengths of about 20 that vary slightly from text to text, and ranks results by raw dot product. What can go wrong?
3. Two unit-length vectors have a cosine similarity of 0.5. What is the Euclidean (L2) distance between them?
4. A RAG bot answers "I do not know" whenever the best match scores below 0.8. It was tuned on mxbai-embed-large, and the team switches to nomic-embed-text. Based on this chapter, what is the likely effect?

What's next​

Every search so far has compared the query against every single article, which is fine for 16. In Chapter 3, you will find out how far that brute-force approach goes, why the obvious speedup, a tree index, fails in high dimensions for exactly the reason you just measured, and where the nearest-neighbor idea came from in the first place.