Lesson 3 of 6

Word embeddings

Give every word a list of numbers, learned from how words are used, and meaning turns into geometry. Explore real word vectors, their neighbors, their analogies, and their biases.

Intermediate18 min

In this lesson you will

  • Explain how word embeddings are learned from the contexts words appear in
  • Find a word's neighbors and solve analogies with vector arithmetic
  • Recognize what embeddings absorb from their training text, including bias

“You shall know a word by the company it keeps,” wrote the linguist J. R. Firth in 1957. Words that show up in similar places tend to mean similar things: coffee and tea both get brewed, poured, and drunk; piano and violin both get played and tuned. That idea, called the distributional hypothesis, is all it takes to turn words into useful vectors.

Learning vectors from text

A word EmbeddingA list of numbers (a vector) that represents a token, word, sentence, or image, learned so that similar things end up with similar vectors.Open in glossary gives each word a vector, typically 50 to 300 numbers long, chosen so that words used in similar contexts get similar vectors. Two landmark methods made this practical. word2vec (Mikolov and colleagues at Google, 2013) trains a small network to predict the words around each word. GloVe (Pennington, Socher, and Manning at Stanford, 2014) fits vectors directly to counts of how often each pair of words appears near each other.

The demos below use real GloVe vectors trained on about six billion words from Wikipedia and newswire text. Each of the 10,000 words here is a list of 100 numbers. Nobody chose those numbers, and none of them means “size” or “color” on its own. Meaning lives in the directions and distances between the vectors: each word becomes a point in a Latent spaceThe space of learned, hidden coordinates a model uses to represent things internally, such as word vectors or a network's hidden layers. Nearby points stand for similar things.Open in glossary, a space of learned coordinates that no one labeled.

Explore real word embeddings

10,000 common English words, each a list of 100 numbers learned from billions of words of text (GloVe).

View

Loading 1.1 MB of word vectors...

Try this

  • Start with piano, then select violin, then orchestra in the results. You are walking through a neighborhood of related words.
  • Try virus and coffee. The neighbors are not dictionary synonyms; they are words that show up in similar sentences.
  • Try happy. Its neighbors include words like “glad” but also “really” and “sure”. Words that share contexts are not always words that share meaning.
  • Type a word that is not in the set, such as a rare name. Every model has a vocabulary, and anything outside it has no vector at all.

Behind every list is the cosine similarity from the previous lesson, computed between one word and the 9,999 others.

Meaning as direction

The most famous property of word vectors is that differences between them can be meaningful. If you take the vector for king, subtract man, and add woman, the nearest word to the result, other than the three you started with, is queen. This vector offset method was introduced with word2vec and is often called 3CosAdd: to complete “a is to b as c is to ?”, compute b−a+c\mathbf{b} - \mathbf{a} + \mathbf{c} and find the nearest word, leaving out aa, bb, and cc themselves.

Explore real word embeddings

10,000 common English words, each a list of 100 numbers learned from billions of words of text (GloVe).

View

Loading 1.1 MB of word vectors...

Try this

  • Run the first five examples. Capitals, family roles, verb forms, and comparatives all work, using nothing but addition and subtraction.
  • Run big : bigger :: small. The intended answer, “smaller”, comes second. Analogies are approximate, and the demo reports what the vectors actually say.
  • Run japan : yen :: europe. “francs” edges out “euro”. The training text includes news from the 1990s, before the euro replaced the French franc.
  • Make up your own, such as paris : france :: berlin or slow : slower :: fast.

Why does this work? If the training text uses king and queen in nearly the same contexts except for the ones that differ by gender, and the same is true for man and woman, then the step from one to the other ends up pointing in a similar direction in both pairs. The map view shows this directly: each arrow goes from one word in a pair to the other.

Explore real word embeddings

10,000 common English words, each a list of 100 numbers learned from billions of words of text (GloVe).

View

Loading 1.1 MB of word vectors...

Try this

  • With Countries, the arrows from each country to its capital run nearly parallel. The readout puts a number on it for this 2D view; in the full 100 dimensions the arrows agree less (about 0.8 instead of nearly 1), because a projection fitted to just these words tends to line them up.
  • Switch to Past tense: also nearly parallel. Then Gender and Comparatives: mostly parallel, with a few arrows that disagree. Real data is messier than the famous examples suggest.
  • Read the Variance kept readout. This flat picture keeps only about 40 percent of the variation among these 20 words, and less of everything else the 100-number vectors encode. The next lesson is about what gets lost.

Embeddings absorb their training text

Embeddings learn whatever patterns are in the text, including ones nobody wants. In the analogies view, man : doctor :: woman returns “nurse”. In 2016, Bolukbasi and colleagues documented this kind of gender stereotyping in word2vec vectors in a paper titled “Man is to Computer Programmer as Woman is to Homemaker?”, and warned that it can carry into systems built on the vectors.

There is a second, subtler lesson in the same example. The standard method never returns an input word, so “doctor” itself cannot be the answer. In 2020 Nissim, van Noord, and van der Goot showed that when that rule is lifted, the original word often comes back first, and that some widely shared biased analogies looked more extreme than the vectors support. Both things are true: the vectors do encode stereotypes from text, and the way results are presented can exaggerate them.

Key ideas

  • Word embeddings are learned from which words appear near which others, following the distributional hypothesis.
  • Similar words have similar vectors, so nearest neighbors by cosine similarity group related words.
  • Some relationships become directions, so vector arithmetic can solve analogies, approximately.
  • Embeddings reflect their training text, including its era and its stereotypes.
  • Static word embeddings give each word one vector; modern models make vectors depend on context.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1Where do the numbers in a word embedding like GloVe come from?
2Computing king - man + woman and finding the nearest word, leaving out the three inputs, gives queen. What does this show?
3The analogy man is to doctor as woman is to ? returns nurse. What is the best reading?

Progress is saved in this browser only.

Up nextSeeing high dimensions
Next
Language as vectors
  1. 1Tokens
  2. 2Vectors and similarity
  3. 3Word embeddings
  4. 4Seeing high dimensions
  5. 5Search by meaning
  6. 6Retrieval-augmented generation

Try "embedding", "softmax", "overfitting", or "backpropagation".