Language as vectors

Models do not read words; they read numbers. Split text into tokens, map words to points in space where meaning becomes direction, and search by meaning with a real embedding model running in your browser.

Level
Intermediate
Length
6 lessons, 1 hr 32 min
Assumes
Comfort with the idea of a vector as a list of numbers.
Builds on
How machines learn
  1. 1TokensLanguage models never see letters or words. They see tokens, chunks of text learned from data. Build a tokenizer step by step, then look at text through a real one.15 min
  2. 2Vectors and similarityAn embedding is a list of numbers, which is to say a point in space. Learn the three standard ways to measure how close two of them are, and when each one is the right choice.12 min
  3. 3Word embeddingsGive every word a list of numbers, learned from how words are used, and meaning turns into geometry. Explore real word vectors, their neighbors, their analogies, and their biases.18 min
  4. 4Seeing high dimensionsEmbeddings live in hundreds of dimensions, and we can only look at two or three. Learn how projection works, what it hides, and why distance itself behaves strangely when dimensions pile up.16 min
  5. 5Search by meaningA sentence embedding model turns a whole sentence into one vector, so you can search by what text means instead of which words it uses. Run a real one in your browser and race it against keyword search.15 min
  6. 6Retrieval-augmented generationLanguage models only know what was in their training data. Retrieval-augmented generation looks up relevant passages first and puts them in the prompt. Build the pipeline step by step on a document no model has seen.16 min

Labs in this track

Try "embedding", "softmax", "overfitting", or "backpropagation".