Every picture of embeddings you have ever seen, including the map in the last lesson, is a shadow. The real vectors have 100, 384, or thousands of numbers. To look at them, we have to choose a few directions and flatten everything onto them. That is useful, and it is also the easiest way to fool yourself.
Projection
The most common way to flatten data is Principal component analysisA method that finds the directions along which data varies the most, often used to project many-dimensional data down to two or three dimensions for viewing.Open in glossary (PCA). It finds the directions along which the data is most spread out:
- Subtract the average vector, so the cloud of points is centered on the origin.
- Find the direction along which the points vary the most. That is the first principal component.
- Find the direction of greatest remaining variation at right angles to the first. That is the second. Repeat for the third.
- Describe each point by how far it lies along each of those directions.
Each component comes with a number: the share of the total variance it captures. Add them up and you know how much of the spread your picture keeps, and how much it hides.
Look inside a 100-dimensional space
Word vectors flattened to their three most informative directions. Drag to turn it. Tap a word to compare its true neighbors with what you see.
Loading 1.1 MB of word vectors...
Try this
- Drag to turn the cloud (on a phone, drag sideways and use the Tilt slider). The animals sit apart from the rest. Food and sports overlap from some angles and separate from others: in 100 dimensions they are distinct, but a flat view can stack them. Nobody labeled any of these groups; they come from how the words are used.
- Read Variance kept. Three directions out of 100 keep about half of the variance among these 40 words. The other half is invisible here.
- tiger is selected. The orange lines go to its three nearest neighbors measured in all 100 dimensions. Compare them with the words nearest to it on screen, and turn the cloud: the on-screen neighbors change with the view, while the true ones never do.
- Turn on all six groups. The variance kept drops to about a third, and more words now sit beside words from other groups on screen, even though their true neighbors are still in their own group.
The math of principal componentsOptional
Stack the centered vectors as the rows of a matrix with rows (one per word) and columns (one per dimension). Their covariance matrix is
a matrix whose entry measures how dimensions and vary together. The principal components are the eigenvectors of , unit vectors with , and each eigenvalue is the variance of the data along its eigenvector. The share of variance kept by the first components is
The demo finds the top eigenvectors by orthogonal iteration: start with random directions, multiply by again and again, and re-orthogonalize each time. Multiplying by stretches every direction in proportion to its variance, so the directions settle onto the components with the largest eigenvalues.
Distance in many dimensions
Projection is not the only surprise. Distance itself changes character as dimensions grow. In two dimensions, some random points are close neighbors and others are far apart. In a thousand dimensions, almost every pair of random points is about the same distance apart.
Distances in many dimensions
160 random points in a cube. The chart shows every pairwise distance, divided by the average distance.
Try this
- At 2 dimensions, the distances spread widely, and a typical point’s farthest neighbor is many times as far away as its nearest.
- Drag Dimensions up to 10, 100, and 1,000. The histogram squeezes into a narrow spike around the average, and the farthest-to-nearest ratio falls toward 1.
This effect, part of what is called the Curse of dimensionalityThe collection of ways data behaves unintuitively in many dimensions. For example, distances between random points become nearly all the same, so nearest neighbors stop standing out.Open in glossary, was analyzed by Beyer and colleagues in 1999 under the title “When Is ‘Nearest Neighbor’ Meaningful?”. Their answer, for points spread at random, was: increasingly, not at all. The reason is that each dimension adds a little independent noise to every distance, and the sum of many independent contributions varies less and less relative to its size.
So why does searching embeddings work? Because learned embeddings are nothing like random points. Training packs related items close together and pushes unrelated ones apart, so the data occupies thin, structured regions of the space rather than filling it evenly. Within that structure the difference between near and far stays large. The demo shows the worst case; real embeddings largely avoid it, though some concentration remains (unrelated texts often still score a cosine of 0.2 or more).
Key ideas
- Any picture of embeddings is a projection that keeps a few directions and flattens the rest.
- PCA keeps the directions of greatest variance and tells you how much variance the picture keeps.
- Check projections against true neighbors: points that look close may not be.
- For random points, distances concentrate as dimensions grow. Learned embeddings largely avoid this because they are highly structured.