Glossary

Plain-language definitions for the terms used across every track, each linked to the lesson that teaches it.

Activation functionalso nonlinearity

The function a neuron applies to its weighted sum, such as the sigmoid, tanh, or ReLU. Without it, stacking layers of neurons would add no power.

Adam

A widely used optimizer that combines momentum with a separate, automatically scaled step size for each parameter. Most large neural networks are trained with Adam or a close variant.

Agentalso AI agent, agentic system

A system that runs a language model in a loop with tools, letting it take several steps, observe results, and decide what to do next toward a goal. Definitions vary in how much autonomy they imply.

Artificial neuronalso neuron, unit, perceptron

The basic unit of a neural network. It multiplies each input by a weight, adds them up with a bias, and passes the total through an activation function.

The idea was loosely inspired by brain cells, but an artificial neuron is a simple formula. Real neurons are far more complicated.

Attention

A mechanism that lets each token build a new vector as a weighted mix of other tokens' value vectors, with weights set by how well its query matches their keys.

Scaled dot-product attention computes softmax(QK⊤/dk)V\mathrm{softmax}(QK^\top/\sqrt{d_k})V.

Attention headalso head, heads

One independent attention computation inside a multi-head attention layer, with its own learned projections and its own pattern of weights.

Autoregressive generationalso autoregressive

Producing a sequence one element at a time, where each new element is predicted from all the ones before it and then appended to the input.

Backpropagationalso backprop, reverse-mode differentiation

The algorithm that computes the gradient of the loss with respect to every weight in a network in one backward pass, by applying the chain rule from the loss back to the inputs.

Base modelalso foundation model, pretrained model

A language model after pretraining only. It continues text in whatever style the input suggests, rather than following instructions or answering like an assistant.

Batch sizealso minibatch, mini-batch

How many training examples are used to compute each gradient step. Small batches give noisy but frequent updates; large batches give smoother but fewer updates per epoch.

Bias (neuron)

A number a neuron adds to its weighted sum before the activation function. It shifts how easily the neuron switches on, regardless of the inputs.

Not to be confused with bias in data, where a dataset misrepresents some situations or groups. See data bias.

BM25also Okapi BM25

A standard keyword ranking formula. It scores documents by the query words they contain, weighting rare words more and limiting the reward for repeated words.

Byte-pair encodingalso BPE

A way to build a subword vocabulary by starting from single characters or bytes and repeatedly merging the most frequent neighboring pair into a new token.

Introduced as a compression method by Philip Gage in 1994 and adapted for subword tokenization by Sennrich, Haddow, and Birch in 2016. Most GPT-style tokenizers use a byte-level variant.

Causal maskalso causal attention, look-ahead mask

A mask that stops each token from attending to tokens after it, by setting those attention scores to negative infinity before the softmax. Used by models that generate text left to right.

Centroid

The average position of a group of points. In k-means, each cluster is represented by its centroid.

Chain of thoughtalso CoT, chain-of-thought prompting

Intermediate reasoning steps a language model writes before its final answer. Prompting for them, or training models to produce them, often improves results on multi-step problems.

Chain rule

The calculus rule for nested functions: if L depends on a and a depends on z, then dL/dz = dL/da times da/dz. Backpropagation applies it over and over.

Chunking

Splitting documents into smaller passages before embedding them, so each passage can be retrieved and fit in a prompt on its own.

Classification

A prediction task where the answer is one of a fixed set of categories, such as spam or not spam, or which digit a drawing shows.

Classifier-free guidancealso CFG, guidance scale

A sampling technique for conditional diffusion models that compares the model's prediction with and without the prompt and moves further in the prompt's direction, trading variety for closer prompt matching.

Computational graph

A diagram of a calculation as simple operations connected by the values flowing between them. Each operation knows its own local derivative, which makes backpropagation mechanical.

Confusion matrix

A table counting a classifier's results by true class and predicted class: true positives, false positives, false negatives, and true negatives.

Context windowalso context length

The maximum number of tokens a language model can take into account at once, covering both the prompt and the text it generates.

Contrastive learningalso InfoNCE, contrastive loss

Training that pulls the embeddings of matching pairs together and pushes non-matching pairs apart, so the model learns a useful space without category labels.

Convolution

Sliding a small grid of weights (a kernel) across an image and computing a weighted sum at every position, producing a new image called a feature map.

Convolutional neural networkalso CNN, ConvNet

A neural network whose early layers are convolutions, so the same small learned patterns are detected everywhere in an image. The standard design for image recognition for over a decade.

Cosine similarity

A measure of how closely two vectors point in the same direction, from -1 (opposite) through 0 (perpendicular) to 1 (same direction), ignoring their lengths.

It is the dot product of the two vectors divided by the product of their lengths, which equals the cosine of the angle between them.

Cross-entropyalso log loss, binary cross-entropy

A loss for predicted probabilities: minus the log of the probability the model gave to the right answer. Confident wrong answers are punished very heavily.

Curse of dimensionality

The collection of ways data behaves unintuitively in many dimensions. For example, distances between random points become nearly all the same, so nearest neighbors stop standing out.

Data augmentation

Making extra training examples by applying small, label-preserving changes to existing ones, such as shifting, rotating, or rescaling an image.

Data biasalso dataset bias, sampling bias

When training data misrepresents the situations or people a model will be used on, for example by leaving some groups out, so the model works worse for them.

Decision boundary

The line or surface where a classifier switches from predicting one class to another.

Deep learning

Machine learning with neural networks that have many layers. It powers modern image recognition, speech recognition, translation, and chatbots.

Derivative

The slope of a function at a point: how much the output changes for a tiny change in the input.

Diffusion modelalso denoising diffusion model, DDPM

A generative model that learns to reverse a gradual noising process, so it can turn pure random noise into new data such as images, step by step.

Direct preference optimizationalso DPO

A way to train a language model on pairs of preferred and rejected responses with a single classification-style loss, without a separate reward model or reinforcement learning.

Distribution shiftalso dataset shift, out-of-distribution data

When the data a model meets in use differs from the data it was trained on. Models often get worse under shift, sometimes without becoming any less confident.

Dot productalso inner product, scalar product

The sum of the products of matching entries of two vectors. It equals the product of their lengths times the cosine of the angle between them.

a⋅b=∑iaibi=∥a∥∥b∥cos⁡θ\mathbf{a} \cdot \mathbf{b} = \sum_i a_i b_i = \lVert \mathbf{a} \rVert \lVert \mathbf{b} \rVert \cos \theta.

Embeddingalso embeddings, embedding vector

A list of numbers (a vector) that represents a token, word, sentence, or image, learned so that similar things end up with similar vectors.

Entropy

A measure of how uncertain a probability distribution is, in bits. Zero means one outcome is certain; k bits is as uncertain as a fair choice among 2 to the power k outcomes.

For probabilities pip_i, entropy is H=−∑ipilog⁡2piH = -\sum_i p_i \log_2 p_i.

Epoch

One full pass through the training data. Training usually runs for many epochs, shuffling the examples each time.

Euclidean distancealso L2 distance

The straight-line distance between two points, the square root of the sum of squared differences between matching entries.

Feature

One measurable property of an example, given to a model as a number, such as a fruit's weight or a pixel's brightness.

Feature mapalso activation map, channel

The output of one convolution kernel applied across a whole image: a grid showing how strongly that kernel's pattern appears at each location.

Feed-forward networkalso MLP, FFN, feed-forward layer

The MLP inside each transformer block. It processes every token's vector on its own, usually widening it about four times, applying a nonlinearity, and projecting back. It holds most of each block's parameters.

Fine-tuningalso supervised fine-tuning, SFT

Continuing to train a pretrained model on a smaller, targeted dataset, such as example conversations, to change its behavior. Supervised fine-tuning uses examples of good responses.

Generalization

How well a model performs on new data it was not trained on. It is the goal of machine learning; doing well on the training data is only a means to it.

Generative AIalso GenAI

Models that produce new content, such as text, images, audio, or code, rather than only labeling or scoring existing content.

Gradient

The list of slopes of a function, one per parameter. It points in the direction in which the function increases fastest, so its opposite points downhill.

Gradient descent

An optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.

Greedy decoding

Generating text by always choosing the single most probable next token. Equivalent to sampling at temperature 0.

Grouped-query attentionalso GQA, multi-query attention, MQA

An attention variant in which groups of query heads share one key head and one value head. It shrinks the KV cache with little loss in quality. Multi-query attention is the extreme case of a single shared key/value head.

Hallucinationalso confabulation

When a language model states something false or invented, such as a fake quote or citation, in the same fluent, confident style as true statements.

Hidden layer

A layer of neurons between a network's inputs and its output. Its values are not given or read directly; the network works them out as intermediate steps.

Hyperparameter

A setting chosen by people rather than learned from data, such as the learning rate, the number of layers, or the strength of regularization.

Induction head

An attention head that, at a token A, looks back for earlier copies of A and attends to the token that followed them, helping the model continue repeated patterns.

k-meansalso k-means clustering

A clustering method that places k centers, assigns every point to its nearest center, moves each center to the average of its points, and repeats until nothing changes.

k-nearest neighborsalso k-NN, KNN

A method that predicts the label of a new point by finding the k most similar labeled examples and taking a vote among them.

Kernel (convolution)also filter

The small grid of weights, often 3 x 3, that a convolution slides across an image. In a convolutional network the kernel values are learned.

KV cachealso key-value cache

Stored keys and values for every token already processed, kept between generation steps so each new token only needs its own keys and values computed. It grows with every token in the context.

Its size is 2×2 \times layers ×\times key/value heads ×\times head size ×\times bytes per number for each token.

Labelalso target

The answer a model should learn to give for an example, such as "apple" or "orange". Labeled examples are what supervised learning learns from.

Language model

A model that assigns probabilities to what comes next in a piece of text. Writing text with one means repeatedly predicting the next word or token and picking one.

Large language modelalso LLM

A neural network, usually a transformer with billions of parameters, trained on large amounts of text to predict the next token. Chatbots are built on them.

Latent spacealso latent, latent representation, latent vector

The space of learned, hidden coordinates a model uses to represent things internally, such as word vectors or a network's hidden layers. Nearby points stand for similar things.

Latent means hidden: present, but not directly observed. A model never stores the word “king” or a picture of a cat as such. It stores coordinates, and what they mean lives in where those points sit relative to one another. Word embeddings, the hidden layers of a neural network, and the compressed images that latent diffusion models such as Stable Diffusion work on (Rombach et al., 2022) are all latent spaces.

This site is named after the idea. Much of what makes modern AI work happens in spaces you cannot see directly, and the demos here are ways of looking into them.

Layeralso dense layer, fully connected layer

A group of neurons that all read the same inputs. A fully connected layer computes Wx + b for a weight matrix W and bias vector b, then applies an activation function.

Layer normalizationalso LayerNorm, RMSNorm, normalization

Rescaling each token's vector to a standard mean and spread, then applying a learned scale (and, in LayerNorm, a shift). Keeps every sublayer's input on a stable scale.

LayerNorm (Ba et al., 2016) subtracts the mean and divides by the standard deviation of the vector’s entries. RMSNorm (Zhang and Sennrich, 2019) only divides by the root mean square.

Learning ratealso step size

The step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.

Linear regression

Predicting a number with a weighted sum of the inputs plus a constant. With one input, it is fitting a straight line.

Linear transformationalso linear map, affine transformation

A map that can only rotate, stretch, shear, or reflect space, so straight lines stay straight and parallel lines stay parallel. Adding a bias also shifts it (an affine map).

Local minimum

A point lower than everything around it but not necessarily the lowest point overall. Gradient descent can settle in one and stop improving.

Logistic regression

A classifier that computes a weighted sum of the inputs and passes it through the sigmoid to get the probability of the positive class. Despite the name, it is used for classification.

Logitalso logits

A raw, unnormalized score a model outputs for one option, such as one token in the vocabulary. Softmax turns a vector of logits into probabilities.

Loss functionalso loss, cost function, objective

A formula that turns a model's mistakes on the training data into a single number. Lower is better, and training tries to make it as low as possible.

Machine learningalso ML

A way of building software where the behavior is learned from examples with known answers, instead of being written as rules by a person.

Matrix multiplicationalso matmul

Combining a matrix with a vector or another matrix by taking dot products of rows with columns. An (m x n) matrix times an n-vector gives an m-vector.

Mean absolute erroralso MAE

The average size of the differences between predictions and true values, ignoring sign. Less sensitive to a few wild outliers than mean squared error.

Mean squared erroralso MSE

The average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.

Mixture of expertsalso MoE, sparse mixture of experts

An architecture with many alternative sub-networks (experts) and a router that runs only a few of them for each token, so a model can store many parameters while using few per token.

MNIST

A classic dataset of 70,000 handwritten digits, each a 28 x 28 grayscale image: 60,000 for training and 10,000 for testing.

Model

The rule a learning algorithm produces from data. Given an input, a model returns a prediction. In neural networks, the model is defined by its parameters.

Momentum

An optimizer that keeps a running sum of past gradients, like a ball gathering speed downhill, so it moves faster along consistent slopes and damps zig-zags.

Multi-head attention

Running several attention heads in parallel, each with its own query, key, and value projections, then concatenating their outputs and mixing them with one more matrix.

n-gram modelalso bigram, trigram, n-gram

A simple language model that predicts the next word by counting what followed the previous one or few words in its training text.

Neural networkalso network, neural net

A model built from layers of artificial neurons, where each layer's outputs become the next layer's inputs. Training adjusts all of the weights together.

Optimizer

The rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.

Outlier

A data point far from the pattern of the rest, from a rare event or a recording error. Some loss functions are pulled hard by outliers.

Overfitting

When a model fits the quirks and noise of its training data so closely that it does worse on new data. Training error keeps falling while test error rises.

Parameteralso weight

A number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.

Pixel

One tiny square of a digital image. A grayscale pixel is stored as one brightness number; a color pixel as three (red, green, and blue).

Position encodingalso positional encoding, positional embedding, position embedding

Information about where each token sits in the sequence, given to a transformer because attention on its own ignores word order.

The original transformer added fixed sine and cosine waves to each token’s embedding. GPT-2 learned one vector per position. Many recent models use rotary embeddings (RoPE) instead.

Precision

Of the items a classifier flagged as positive, the fraction that really are positive. High precision means few false alarms.

Pretrainingalso pre-training

The first and most expensive stage of training a language model, in which it learns to predict the next token across a very large amount of text. The result is a base model.

Principal component analysisalso PCA

A method that finds the directions along which data varies the most, often used to project many-dimensional data down to two or three dimensions for viewing.

The principal components are the eigenvectors of the data’s covariance matrix; each eigenvalue is the variance along its component.

Prompt injectionalso indirect prompt injection

An attack in which text the model reads, such as a web page, email, or document, contains instructions that override or redirect what the user or developer intended.

Query, key, and valuealso queries, keys, values, Q K V

The three vectors attention makes from each token. A token's query is compared with other tokens' keys to decide how much of each of their values to take.

Recallalso sensitivity, true positive rate

Of all the items that really are positive, the fraction the classifier flagged. High recall means few misses.

Regression

A prediction task where the answer is a number, such as a price, a temperature, or a time.

Regularizationalso L2 regularization, weight decay, ridge

Any technique that discourages a model from fitting noise, for example adding a penalty on large parameter values to the loss.

ReLUalso rectified linear unit

The rectified linear unit, max(0, z): it passes positive inputs through unchanged and turns negative inputs into zero. The most common activation in deep networks.

Residual streamalso residual connection, skip connection

The running vector for each token that flows through a transformer from the embedding to the output. Every attention and MLP sublayer reads from it and adds its result back.

A residual connection computes x+f(x)x + f(x) instead of f(x)f(x). The name “residual stream” for the sum that accumulates through a transformer was popularized by Elhage et al. (2021).

Retrieval-augmented generationalso RAG

A technique that searches a collection of documents for passages relevant to a question and adds them to a language model's prompt, so it can answer from that text.

RLHFalso reinforcement learning from human feedback

Reinforcement learning from human feedback. People compare pairs of model responses, a reward model learns to predict their preferences, and the language model is then trained with reinforcement learning to score well on it.

ROC curvealso AUC, area under the curve

A plot of the true positive rate against the false positive rate as a classifier's threshold sweeps from strict to lenient. The area under it (AUC) summarizes how well the scores rank positives above negatives.

Rotary position embeddingalso RoPE, rotary embedding

A way to encode position by rotating each pair of query and key dimensions by an angle proportional to the token's position, so attention scores depend on position only through the offset between tokens.

Routeralso gating network, gate

The small learned layer in a mixture-of-experts model that scores every expert for each token and picks which few experts will process it.

Saddle point

A flat point that curves up in some directions and down in others, like the middle of a horse saddle. The gradient is zero there, but it is not a minimum.

Scaling lawsalso scaling law, Chinchilla scaling

Fitted curves that predict how a model's loss falls as parameters, training data, or compute grow. They are empirical, and their constants depend on the architecture, data, and training recipe.

Scorealso score function

The gradient of the log of a probability density, which points in the direction where data becomes more likely. Diffusion models learn the score of noisy data at every noise level.

Self-attention

Attention in which the queries, keys, and values all come from the same sequence, so every token can draw information from every other token in it.

Self-consistencyalso majority voting

Sampling several reasoning paths for the same question and returning the most common final answer. It helps when the correct answer is the model's most likely answer.

Semantic searchalso vector search, embedding search

Search that ranks documents by how similar their meaning is to the query, usually by comparing sentence embeddings with cosine similarity.

Sentence embeddingalso text embedding

A single vector representing a whole sentence or passage, produced by a model trained so that texts with similar meaning get similar vectors.

Sigmoidalso logistic function

The S-shaped function 1 / (1 + e^-z), which squashes any number into the range 0 to 1 so it can be read as a probability.

Softmax

A function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.

softmax(z)i=ezi/∑jezj\mathrm{softmax}(z)_i = e^{z_i} / \sum_j e^{z_j}. Adding the same constant to every score leaves the result unchanged.

Stochastic gradient descentalso SGD, mini-batch gradient descent

Gradient descent where each step uses a small random batch of training examples instead of all of them. Steps are noisier but far cheaper, which is how large models are trained.

Supervised learning

Learning from examples that come with the right answer attached, such as photos labeled "cat" or houses with their sale prices.

Tanh

The hyperbolic tangent, an S-shaped activation function that squashes any number into the range from -1 to 1. It is a rescaled sigmoid centered on zero.

Temperature

A setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.

Test setalso test data, held-out data

Examples held back from training and used only to measure how well a model does on data it has never seen.

Test-time computealso inference-time compute

Computation a model spends while answering, rather than during training, such as writing out reasoning, sampling several answers, or searching over candidates. More of it often improves accuracy on hard problems.

Token

The unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.

Tokenizer

The program that splits text into tokens and maps each one to an id in a vocabulary, and turns ids back into text.

Tool usealso function calling, tool calling

Letting a language model request actions from external programs, such as searches, calculations, or API calls. The model outputs a structured call; the application runs it and returns the result.

Top-k samplingalso top-k

Sampling the next token from only the k most probable tokens, after setting all others to zero probability and renormalizing.

Top-p samplingalso nucleus sampling, top-p

Sampling the next token from the smallest set of most probable tokens whose probabilities add up to at least p. Also called nucleus sampling.

Training dataalso training set

The examples, usually with known answers, that a model learns from. A model can only learn patterns that are present in its training data.

Transformer

A neural network architecture built from stacked layers of attention and feed-forward networks, introduced by Vaswani et al. in 2017. Most modern language models are transformers.

Underfitting

When a model is too simple to capture the real pattern, so it does poorly on both the training data and new data.

Unsupervised learning

Learning from examples that have no answers attached, by finding structure such as groups or directions of variation in the data itself.

Validation setalso validation data, dev set

A second held-back set used while building a model to choose settings such as the learning rate or model size, so the test set stays untouched until the final check.

Vector

An ordered list of numbers. It can also be read as an arrow, or a point, in a space with one dimension per number.

Vector database

A database built to store embeddings and quickly find the ones most similar to a query vector, usually with an approximate nearest neighbor index.

Vocabularyalso vocab

The fixed list of tokens a tokenizer and model know. Each token's position in the list is its id.

Weight

A number that sets how much one input matters to a neuron, and in which direction. Positive weights push the output up, negative weights push it down.

Zero-shotalso zero-shot learning, zero-shot classification

Doing a task without any examples of that specific task during training or in the prompt, for example classifying images into categories the model was never trained to label.

Try "embedding", "softmax", "overfitting", or "backpropagation".