Lesson 2 of 6

Vectors and similarity

An embedding is a list of numbers, which is to say a point in space. Learn the three standard ways to measure how close two of them are, and when each one is the right choice.

Intermediate12 min

In this lesson you will

  • Read a vector as both a list of numbers and an arrow in space
  • Compute a dot product, cosine similarity, and Euclidean distance
  • Explain why cosine similarity ignores length, and when that matters

In the last lesson, text became a sequence of token ids. Ids are just labels, though: token 9,067 is not “closer” to token 9,068 in any useful sense. To compute with meaning, a model turns each token, word, or whole sentence into a VectorAn ordered list of numbers. It can also be read as an arrow, or a point, in a space with one dimension per number.Open in glossary, a list of numbers. Once things are vectors, “how similar are these?” becomes a geometry question with exact answers.

A list of numbers is a point

The vector (3,1)(3, 1) is two numbers. It is also an arrow from the origin to the point three steps right and one step up. Both readings are always true, and switching between them is the key skill of this track.

Real embeddings have hundreds or thousands of numbers. You cannot draw an arrow in 384 dimensions, but every formula below works the same way in any number of dimensions. That is why it is worth getting them into your fingers in two dimensions first.

Three ways to compare two vectors

Dot product, cosine, and distance

Drag the tips of a and b. Three ways to compare two vectors update as you move them.

Dot product a · b6.50
Cosine similarity0.82
Angle35°
Distance ‖a − b‖1.80

Try this

  • Choose Same direction. The angle is 0 and the cosine similarity is exactly 1, even though b is three times as long. Now look at the dot product and the distance: both care about length.
  • Choose Perpendicular. The dot product is 0. Drag b a little toward a and watch it turn positive.
  • Choose Opposite. The cosine is close to -1: the vectors point nearly opposite ways.
  • Turn on Scale both to length 1. Now both arrows end on the dashed circle, the dot product equals the cosine, and a smaller distance always means a larger cosine.
  • Focus a handle with Tab and move it with the arrow keys, holding Shift for bigger steps.

Dot product

The Dot productThe sum of the products of matching entries of two vectors. It equals the product of their lengths times the cosine of the angle between them.Open in glossary multiplies matching entries and adds them up:

a⋅b=∑i=1naibi=a1b1+a2b2+⋯+anbn\mathbf{a} \cdot \mathbf{b} = \sum_{i=1}^{n} a_i b_i = a_1 b_1 + a_2 b_2 + \dots + a_n b_n

Here a\mathbf{a} and b\mathbf{b} are vectors with nn entries each, and aia_i is the ii-th entry of a\mathbf{a}. The sum is large and positive when the two vectors have big entries in the same places with the same signs. It is the workhorse of neural networks: every layer computes many dot products, and attention, which you will meet in the transformers track, is built on them.

Geometrically, the dot product is the product of the two lengths and the cosine of the angle θ\theta between them:

a⋅b=∥a∥ ∥b∥cos⁡θ\mathbf{a} \cdot \mathbf{b} = \lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert \cos \theta

where ∥a∥=a12+⋯+an2\lVert \mathbf{a} \rVert = \sqrt{a_1^2 + \dots + a_n^2} is the length of a\mathbf{a}.

Cosine similarity

Divide out the lengths and only the angle is left. That is Cosine similarityA measure of how closely two vectors point in the same direction, from -1 (opposite) through 0 (perpendicular) to 1 (same direction), ignoring their lengths.Open in glossary:

cos⁡θ=a⋅b∥a∥ ∥b∥\cos \theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert}

It runs from 11 (same direction) through 00 (perpendicular) to −1-1 (opposite). Because it ignores length, it answers “are these about the same thing?” without being swayed by how strongly each one says it. That makes it the default way to compare embeddings.

Euclidean distance

Euclidean distanceThe straight-line distance between two points, the square root of the sum of squared differences between matching entries.Open in glossary is the straight-line distance between the two tips, the dashed line in the demo:

∥a−b∥=∑i=1n(ai−bi)2\lVert \mathbf{a} - \mathbf{b} \rVert = \sqrt{\sum_{i=1}^{n} (a_i - b_i)^2}

It is the measure you used in the first lesson of the site, when k-nearest neighbors looked for the closest examples.

When vectors have length 1

Many embedding models scale every output vector to length 1. Then the three measures stop disagreeing. The dot product is the cosine similarity, and the distance is tied to the cosine exactly:

∥a−b∥2=∥a∥2+∥b∥2−2 a⋅b=2−2cos⁡θ\lVert \mathbf{a} - \mathbf{b} \rVert^2 = \lVert \mathbf{a} \rVert^2 + \lVert \mathbf{b} \rVert^2 - 2\,\mathbf{a} \cdot \mathbf{b} = 2 - 2 \cos \theta

So sorting by smallest distance, largest cosine, or largest dot product gives the same order. Search systems take advantage of this: they normalize once, then use the dot product, the cheapest of the three.

Why length can matterOptional

Ignoring length is usually what you want, but not always. In some word embeddings, the length of a word’s vector loosely tracks how often and how consistently the word is used, and some recommendation systems deliberately use the raw dot product so that more popular items score higher. Whether length carries signal depends on how the vectors were trained. Choosing a similarity measure is a modeling decision, not a formality.

Key ideas

  • A vector is a list of numbers and, equally, an arrow from the origin.
  • The dot product multiplies matching entries and adds them. It grows with both alignment and length.
  • Cosine similarity keeps only the angle: 1 same direction, 0 perpendicular, -1 opposite.
  • For vectors of length 1, dot product, cosine, and distance all rank neighbors the same way.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Vector a is (1, 2). Vector b is (3, 6), three times as long and pointing the same way. What is their cosine similarity?
2Two vectors are perpendicular. Which statement is true?
3Many embedding models scale every vector to length 1. Why is that convenient?

Progress is saved in this browser only.

Up nextWord embeddings
Next
Language as vectors
  1. 1Tokens
  2. 2Vectors and similarity
  3. 3Word embeddings
  4. 4Seeing high dimensions
  5. 5Search by meaning
  6. 6Retrieval-augmented generation

Try "embedding", "softmax", "overfitting", or "backpropagation".