The function a neuron applies to its weighted sum, such as the sigmoid, tanh, or ReLU. Without it, stacking layers of neurons would add no power.
Taught in The artificial neuronRelated: Artificial neuron, Sigmoid
Plain-language definitions for the terms used across every track, each linked to the lesson that teaches it.
The function a neuron applies to its weighted sum, such as the sigmoid, tanh, or ReLU. Without it, stacking layers of neurons would add no power.
Taught in The artificial neuronRelated: Artificial neuron, Sigmoid
A widely used optimizer that combines momentum with a separate, automatically scaled step size for each parameter. Most large neural networks are trained with Adam or a close variant.
Taught in Better optimizersRelated: Optimizer, Momentum
A system that runs a language model in a loop with tools, letting it take several steps, observe results, and decide what to do next toward a goal. Definitions vary in how much autonomy they imply.
Taught in Models that use toolsRelated: Tool use, Prompt injection
The basic unit of a neural network. It multiplies each input by a weight, adds them up with a bias, and passes the total through an activation function.
The idea was loosely inspired by brain cells, but an artificial neuron is a simple formula. Real neurons are far more complicated.
Taught in The artificial neuronRelated: Weight, Bias (neuron), Activation function, Neural network
A mechanism that lets each token build a new vector as a weighted mix of other tokens' value vectors, with weights set by how well its query matches their keys.
Scaled dot-product attention computes .
Taught in Attention, step by stepRelated: Self-attention, Query, key, and value, Transformer, Softmax
One independent attention computation inside a multi-head attention layer, with its own learned projections and its own pattern of weights.
Taught in Many heads and the causal maskRelated: Multi-head attention, Induction head
Producing a sequence one element at a time, where each new element is predicted from all the ones before it and then appended to the input.
Taught in Predicting the next tokenRelated: Large language model, Token, Causal mask
The algorithm that computes the gradient of the loss with respect to every weight in a network in one backward pass, by applying the chain rule from the loss back to the inputs.
Taught in BackpropagationRelated: Chain rule, Computational graph, Gradient
A language model after pretraining only. It continues text in whatever style the input suggests, rather than following instructions or answering like an assistant.
Taught in Inside a real language modelRelated: Pretraining, Fine-tuning
How many training examples are used to compute each gradient step. Small batches give noisy but frequent updates; large batches give smoother but fewer updates per epoch.
Taught in Training a networkRelated: Epoch, Stochastic gradient descent
A number a neuron adds to its weighted sum before the activation function. It shifts how easily the neuron switches on, regardless of the inputs.
Not to be confused with bias in data, where a dataset misrepresents some situations or groups. See data bias.
Taught in The artificial neuronRelated: Artificial neuron, Weight, Parameter
A standard keyword ranking formula. It scores documents by the query words they contain, weighting rare words more and limiting the reward for repeated words.
Taught in Search by meaningRelated: Semantic search
A way to build a subword vocabulary by starting from single characters or bytes and repeatedly merging the most frequent neighboring pair into a new token.
Introduced as a compression method by Philip Gage in 1994 and adapted for subword tokenization by Sennrich, Haddow, and Birch in 2016. Most GPT-style tokenizers use a byte-level variant.
Taught in TokensRelated: Tokenizer, Vocabulary, Token
A mask that stops each token from attending to tokens after it, by setting those attention scores to negative infinity before the softmax. Used by models that generate text left to right.
Taught in Many heads and the causal maskRelated: Self-attention, Autoregressive generation
The average position of a group of points. In k-means, each cluster is represented by its centroid.
Taught in Finding groups: k-meansRelated: k-means
Intermediate reasoning steps a language model writes before its final answer. Prompting for them, or training models to produce them, often improves results on multi-step problems.
Taught in Thinking longerRelated: Test-time compute, Self-consistency
The calculus rule for nested functions: if L depends on a and a depends on z, then dL/dz = dL/da times da/dz. Backpropagation applies it over and over.
Taught in BackpropagationRelated: Backpropagation, Derivative
Splitting documents into smaller passages before embedding them, so each passage can be retrieved and fit in a prompt on its own.
Taught in Retrieval-augmented generationRelated: Retrieval-augmented generation, Sentence embedding
A prediction task where the answer is one of a fixed set of categories, such as spam or not spam, or which digit a drawing shows.
Taught in Classification and probabilityRelated: Regression, Logistic regression, Decision boundary
A sampling technique for conditional diffusion models that compares the model's prediction with and without the prompt and moves further in the prompt's direction, trading variety for closer prompt matching.
Taught in Diffusion: from noise to dataRelated: Diffusion model
A diagram of a calculation as simple operations connected by the values flowing between them. Each operation knows its own local derivative, which makes backpropagation mechanical.
Taught in BackpropagationRelated: Backpropagation, Chain rule
A table counting a classifier's results by true class and predicted class: true positives, false positives, false negatives, and true negatives.
Taught in Measuring a classifierRelated: Precision, Recall
The maximum number of tokens a language model can take into account at once, covering both the prompt and the text it generates.
Related: Token, Large language model, Retrieval-augmented generation
Training that pulls the embeddings of matching pairs together and pushes non-matching pairs apart, so the model learns a useful space without category labels.
Taught in One space for images and textRelated: Embedding, Cosine similarity, Zero-shot
Sliding a small grid of weights (a kernel) across an image and computing a weighted sum at every position, producing a new image called a feature map.
Taught in ConvolutionsRelated: Kernel (convolution), Feature map, Convolutional neural network
A neural network whose early layers are convolutions, so the same small learned patterns are detected everywhere in an image. The standard design for image recognition for over a decade.
Taught in ConvolutionsRelated: Convolution, Kernel (convolution), Feature map
A measure of how closely two vectors point in the same direction, from -1 (opposite) through 0 (perpendicular) to 1 (same direction), ignoring their lengths.
It is the dot product of the two vectors divided by the product of their lengths, which equals the cosine of the angle between them.
Taught in Vectors and similarityRelated: Embedding
A loss for predicted probabilities: minus the log of the probability the model gave to the right answer. Confident wrong answers are punished very heavily.
Taught in Classification and probabilityRelated: Loss function, Logistic regression, Softmax
The collection of ways data behaves unintuitively in many dimensions. For example, distances between random points become nearly all the same, so nearest neighbors stop standing out.
Taught in Seeing high dimensionsRelated: Principal component analysis, Euclidean distance, k-nearest neighbors
Making extra training examples by applying small, label-preserving changes to existing ones, such as shifting, rotating, or rescaling an image.
Taught in Reading handwritingRelated: Training data, Overfitting
When training data misrepresents the situations or people a model will be used on, for example by leaving some groups out, so the model works worse for them.
Taught in What AI gets wrongRelated: Training data, Distribution shift
The line or surface where a classifier switches from predicting one class to another.
Taught in Classification and probabilityRelated: Classification, Logistic regression
Machine learning with neural networks that have many layers. It powers modern image recognition, speech recognition, translation, and chatbots.
Taught in What is AI, really?Related: Machine learning, Neural network
The slope of a function at a point: how much the output changes for a tiny change in the input.
Taught in Gradient descentRelated: Gradient
A generative model that learns to reverse a gradual noising process, so it can turn pure random noise into new data such as images, step by step.
Taught in Diffusion: from noise to dataRelated: Score, Classifier-free guidance, Generative AI
A way to train a language model on pairs of preferred and rejected responses with a single classification-style loss, without a separate reward model or reinforcement learning.
Taught in How LLMs are trainedRelated: RLHF, Fine-tuning
When the data a model meets in use differs from the data it was trained on. Models often get worse under shift, sometimes without becoming any less confident.
Taught in What AI gets wrongRelated: Training data, Generalization, Data bias
The sum of the products of matching entries of two vectors. It equals the product of their lengths times the cosine of the angle between them.
.
Taught in Vectors and similarityRelated: Vector, Cosine similarity
A list of numbers (a vector) that represents a token, word, sentence, or image, learned so that similar things end up with similar vectors.
Taught in Word embeddingsRelated: Token, Attention
A measure of how uncertain a probability distribution is, in bits. Zero means one outcome is certain; k bits is as uncertain as a fair choice among 2 to the power k outcomes.
For probabilities , entropy is .
Taught in Inside a real language modelRelated: Softmax, Temperature, Cross-entropy
One full pass through the training data. Training usually runs for many epochs, shuffling the examples each time.
Taught in Training a networkRelated: Batch size, Stochastic gradient descent
The straight-line distance between two points, the square root of the sum of squared differences between matching entries.
Taught in Vectors and similarityRelated: Vector, Cosine similarity, k-nearest neighbors
One measurable property of an example, given to a model as a number, such as a fruit's weight or a pixel's brightness.
Taught in Turning things into numbersRelated: Label, Training data, Model
The output of one convolution kernel applied across a whole image: a grid showing how strongly that kernel's pattern appears at each location.
Taught in ConvolutionsRelated: Convolution, Kernel (convolution)
The MLP inside each transformer block. It processes every token's vector on its own, usually widening it about four times, applying a nonlinearity, and projecting back. It holds most of each block's parameters.
Taught in The transformer blockRelated: Transformer, Residual stream, Attention
Continuing to train a pretrained model on a smaller, targeted dataset, such as example conversations, to change its behavior. Supervised fine-tuning uses examples of good responses.
Taught in How LLMs are trainedRelated: Pretraining, RLHF, Base model
How well a model performs on new data it was not trained on. It is the goal of machine learning; doing well on the training data is only a means to it.
Taught in Overfitting and generalizationRelated: Overfitting, Test set
Models that produce new content, such as text, images, audio, or code, rather than only labeling or scoring existing content.
Taught in What is AI, really?Related: Deep learning, Large language model
The list of slopes of a function, one per parameter. It points in the direction in which the function increases fastest, so its opposite points downhill.
Taught in Gradient descentRelated: Gradient descent, Derivative, Loss function
An optimization method that repeatedly moves the parameters a small step in the direction that lowers the loss fastest, the opposite of the gradient.
Taught in Gradient descentRelated: Gradient, Learning rate, Optimizer, Stochastic gradient descent
Generating text by always choosing the single most probable next token. Equivalent to sampling at temperature 0.
Taught in Predicting the next tokenRelated: Temperature, Top-k sampling, Top-p sampling
An attention variant in which groups of query heads share one key head and one value head. It shrinks the KV cache with little loss in quality. Multi-query attention is the extreme case of a single shared key/value head.
Taught in Generating fast: the KV cacheRelated: KV cache, Multi-head attention, Attention
When a language model states something false or invented, such as a fake quote or citation, in the same fluent, confident style as true statements.
Taught in What AI gets wrongRelated: Language model, Large language model
A setting chosen by people rather than learned from data, such as the learning rate, the number of layers, or the strength of regularization.
Taught in Overfitting and generalizationRelated: Parameter, Validation set
An attention head that, at a token A, looks back for earlier copies of A and attends to the token that followed them, helping the model continue repeated patterns.
Taught in Many heads and the causal maskRelated: Attention head, Multi-head attention
A clustering method that places k centers, assigns every point to its nearest center, moves each center to the average of its points, and repeats until nothing changes.
Taught in Finding groups: k-meansRelated: Unsupervised learning, Centroid
A method that predicts the label of a new point by finding the k most similar labeled examples and taking a vote among them.
Taught in What is AI, really?Related: Machine learning, Classification
The small grid of weights, often 3 x 3, that a convolution slides across an image. In a convolutional network the kernel values are learned.
Taught in ConvolutionsRelated: Convolution, Feature map
Stored keys and values for every token already processed, kept between generation steps so each new token only needs its own keys and values computed. It grows with every token in the context.
Its size is layers key/value heads head size bytes per number for each token.
Taught in Generating fast: the KV cacheRelated: Attention, Grouped-query attention, Autoregressive generation
The answer a model should learn to give for an example, such as "apple" or "orange". Labeled examples are what supervised learning learns from.
Taught in Turning things into numbersRelated: Feature, Training data, Supervised learning
A model that assigns probabilities to what comes next in a piece of text. Writing text with one means repeatedly predicting the next word or token and picking one.
Taught in How a chatbot writesRelated: Large language model, n-gram model, Token, Temperature
A neural network, usually a transformer with billions of parameters, trained on large amounts of text to predict the next token. Chatbots are built on them.
Taught in Predicting the next tokenRelated: Token, Transformer, Autoregressive generation
The space of learned, hidden coordinates a model uses to represent things internally, such as word vectors or a network's hidden layers. Nearby points stand for similar things.
Latent means hidden: present, but not directly observed. A model never stores the word “king” or a picture of a cat as such. It stores coordinates, and what they mean lives in where those points sit relative to one another. Word embeddings, the hidden layers of a neural network, and the compressed images that latent diffusion models such as Stable Diffusion work on (Rombach et al., 2022) are all latent spaces.
This site is named after the idea. Much of what makes modern AI work happens in spaces you cannot see directly, and the demos here are ways of looking into them.
Taught in Word embeddingsRelated: Embedding, Vector, Principal component analysis, Diffusion model
A group of neurons that all read the same inputs. A fully connected layer computes Wx + b for a weight matrix W and bias vector b, then applies an activation function.
Taught in Layers are matrix multiplicationsRelated: Neural network, Hidden layer, Matrix multiplication
Rescaling each token's vector to a standard mean and spread, then applying a learned scale (and, in LayerNorm, a shift). Keeps every sublayer's input on a stable scale.
LayerNorm (Ba et al., 2016) subtracts the mean and divides by the standard deviation of the vector’s entries. RMSNorm (Zhang and Sennrich, 2019) only divides by the root mean square.
Taught in The transformer blockRelated: Residual stream, Transformer
The step size in gradient descent. Too small and training crawls; too large and each step overshoots, so the loss can grow instead of shrink.
Taught in Gradient descentRelated: Gradient descent, Optimizer
Predicting a number with a weighted sum of the inputs plus a constant. With one input, it is fitting a straight line.
Taught in Linear regression and lossRelated: Regression, Mean squared error
A map that can only rotate, stretch, shear, or reflect space, so straight lines stay straight and parallel lines stay parallel. Adding a bias also shifts it (an affine map).
Taught in Why nonlinearity mattersRelated: Layer, Activation function
A point lower than everything around it but not necessarily the lowest point overall. Gradient descent can settle in one and stop improving.
Taught in Gradient descentRelated: Gradient descent, Saddle point
A classifier that computes a weighted sum of the inputs and passes it through the sigmoid to get the probability of the positive class. Despite the name, it is used for classification.
Taught in Classification and probabilityRelated: Sigmoid, Cross-entropy, Decision boundary
A raw, unnormalized score a model outputs for one option, such as one token in the vocabulary. Softmax turns a vector of logits into probabilities.
Taught in Predicting the next tokenRelated: Softmax, Temperature
A formula that turns a model's mistakes on the training data into a single number. Lower is better, and training tries to make it as low as possible.
Taught in The learning loopRelated: Mean squared error, Cross-entropy, Gradient descent
A way of building software where the behavior is learned from examples with known answers, instead of being written as rules by a person.
Taught in What is AI, really?Related: Model, Deep learning, Training data
Combining a matrix with a vector or another matrix by taking dot products of rows with columns. An (m x n) matrix times an n-vector gives an m-vector.
Taught in Layers are matrix multiplicationsRelated: Layer, Parameter
The average size of the differences between predictions and true values, ignoring sign. Less sensitive to a few wild outliers than mean squared error.
Taught in Linear regression and lossRelated: Mean squared error, Outlier
The average of the squared differences between predictions and true values. Squaring makes big misses count much more than small ones.
Taught in Linear regression and lossRelated: Mean absolute error, Loss function, Linear regression
An architecture with many alternative sub-networks (experts) and a router that runs only a few of them for each token, so a model can store many parameters while using few per token.
Taught in Mixture of expertsRelated: Router, Parameter
A classic dataset of 70,000 handwritten digits, each a 28 x 28 grayscale image: 60,000 for training and 10,000 for testing.
Taught in Reading handwritingRelated: Training data, Test set
The rule a learning algorithm produces from data. Given an input, a model returns a prediction. In neural networks, the model is defined by its parameters.
Taught in What is AI, really?Related: Machine learning, Parameter
An optimizer that keeps a running sum of past gradients, like a ball gathering speed downhill, so it moves faster along consistent slopes and damps zig-zags.
Taught in Better optimizersRelated: Optimizer, Gradient descent, Adam
Running several attention heads in parallel, each with its own query, key, and value projections, then concatenating their outputs and mixing them with one more matrix.
Taught in Many heads and the causal maskRelated: Attention, Attention head, Transformer
A simple language model that predicts the next word by counting what followed the previous one or few words in its training text.
Taught in How a chatbot writesRelated: Language model
A model built from layers of artificial neurons, where each layer's outputs become the next layer's inputs. Training adjusts all of the weights together.
Taught in Networks of neuronsRelated: Artificial neuron, Hidden layer, Deep learning
The rule that decides how to change a model's parameters at each training step, given the gradient. Gradient descent, momentum, and Adam are optimizers.
Taught in Better optimizersRelated: Gradient descent, Momentum, Adam
A data point far from the pattern of the rest, from a rare event or a recording error. Some loss functions are pulled hard by outliers.
Taught in Linear regression and lossRelated: Mean squared error, Mean absolute error
When a model fits the quirks and noise of its training data so closely that it does worse on new data. Training error keeps falling while test error rises.
Taught in Overfitting and generalizationRelated: Underfitting, Generalization, Regularization, Test set
A number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.
Taught in The learning loopRelated: Model, Gradient descent
One tiny square of a digital image. A grayscale pixel is stored as one brightness number; a color pixel as three (red, green, and blue).
Taught in Turning things into numbersRelated: Feature
Information about where each token sits in the sequence, given to a transformer because attention on its own ignores word order.
The original transformer added fixed sine and cosine waves to each token’s embedding. GPT-2 learned one vector per position. Many recent models use rotary embeddings (RoPE) instead.
Taught in Where words areRelated: Rotary position embedding, Self-attention, Transformer
Of the items a classifier flagged as positive, the fraction that really are positive. High precision means few false alarms.
Taught in Measuring a classifierRelated: Recall, Confusion matrix
The first and most expensive stage of training a language model, in which it learns to predict the next token across a very large amount of text. The result is a base model.
Taught in How LLMs are trainedRelated: Base model, Fine-tuning, Large language model
A method that finds the directions along which data varies the most, often used to project many-dimensional data down to two or three dimensions for viewing.
The principal components are the eigenvectors of the data’s covariance matrix; each eigenvalue is the variance along its component.
Taught in Seeing high dimensionsRelated: Curse of dimensionality, Embedding
An attack in which text the model reads, such as a web page, email, or document, contains instructions that override or redirect what the user or developer intended.
Taught in Models that use toolsRelated: Agent, Tool use
The three vectors attention makes from each token. A token's query is compared with other tokens' keys to decide how much of each of their values to take.
Taught in Attention, step by stepRelated: Attention, Self-attention
Of all the items that really are positive, the fraction the classifier flagged. High recall means few misses.
Taught in Measuring a classifierRelated: Precision, Confusion matrix, ROC curve
A prediction task where the answer is a number, such as a price, a temperature, or a time.
Taught in The learning loopRelated: Classification, Linear regression, Supervised learning
Any technique that discourages a model from fitting noise, for example adding a penalty on large parameter values to the loss.
Taught in Overfitting and generalizationRelated: Overfitting, Hyperparameter
The rectified linear unit, max(0, z): it passes positive inputs through unchanged and turns negative inputs into zero. The most common activation in deep networks.
Taught in Why nonlinearity mattersRelated: Activation function, Tanh, Sigmoid
The running vector for each token that flows through a transformer from the embedding to the output. Every attention and MLP sublayer reads from it and adds its result back.
A residual connection computes instead of . The name “residual stream” for the sum that accumulates through a transformer was popularized by Elhage et al. (2021).
Taught in The transformer blockRelated: Transformer, Layer normalization, Feed-forward network
A technique that searches a collection of documents for passages relevant to a question and adds them to a language model's prompt, so it can answer from that text.
Taught in Retrieval-augmented generationRelated: Semantic search, Chunking, Context window, Hallucination
Reinforcement learning from human feedback. People compare pairs of model responses, a reward model learns to predict their preferences, and the language model is then trained with reinforcement learning to score well on it.
Taught in How LLMs are trainedRelated: Fine-tuning, Direct preference optimization
A plot of the true positive rate against the false positive rate as a classifier's threshold sweeps from strict to lenient. The area under it (AUC) summarizes how well the scores rank positives above negatives.
Taught in Measuring a classifierRelated: Recall, Precision, Confusion matrix
A way to encode position by rotating each pair of query and key dimensions by an angle proportional to the token's position, so attention scores depend on position only through the offset between tokens.
Taught in Where words areRelated: Position encoding, Attention
The small learned layer in a mixture-of-experts model that scores every expert for each token and picks which few experts will process it.
Taught in Mixture of expertsRelated: Mixture of experts
A flat point that curves up in some directions and down in others, like the middle of a horse saddle. The gradient is zero there, but it is not a minimum.
Taught in Better optimizersRelated: Local minimum, Gradient descent
Fitted curves that predict how a model's loss falls as parameters, training data, or compute grow. They are empirical, and their constants depend on the architecture, data, and training recipe.
Taught in How LLMs are trainedRelated: Pretraining, Parameter
The gradient of the log of a probability density, which points in the direction where data becomes more likely. Diffusion models learn the score of noisy data at every noise level.
Taught in Diffusion: from noise to dataRelated: Diffusion model
Attention in which the queries, keys, and values all come from the same sequence, so every token can draw information from every other token in it.
Taught in Attention, step by stepRelated: Attention, Causal mask, Query, key, and value
Sampling several reasoning paths for the same question and returning the most common final answer. It helps when the correct answer is the model's most likely answer.
Taught in Thinking longerRelated: Chain of thought, Test-time compute
Search that ranks documents by how similar their meaning is to the query, usually by comparing sentence embeddings with cosine similarity.
Taught in Search by meaningRelated: Sentence embedding, BM25, Vector database
A single vector representing a whole sentence or passage, produced by a model trained so that texts with similar meaning get similar vectors.
Taught in Search by meaningRelated: Embedding, Semantic search, Contrastive learning
The S-shaped function 1 / (1 + e^-z), which squashes any number into the range 0 to 1 so it can be read as a probability.
Taught in Classification and probabilityRelated: Logistic regression, Logit
A function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.
. Adding the same constant to every score leaves the result unchanged.
Taught in Predicting the next tokenRelated: Logit, Temperature, Attention
Gradient descent where each step uses a small random batch of training examples instead of all of them. Steps are noisier but far cheaper, which is how large models are trained.
Taught in The learning loopRelated: Gradient descent, Batch size
Learning from examples that come with the right answer attached, such as photos labeled "cat" or houses with their sale prices.
Taught in The learning loopRelated: Unsupervised learning, Regression, Classification, Training data
The hyperbolic tangent, an S-shaped activation function that squashes any number into the range from -1 to 1. It is a rescaled sigmoid centered on zero.
Taught in Why nonlinearity mattersRelated: Activation function, Sigmoid, ReLU
A setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.
Taught in Predicting the next tokenRelated: Softmax, Logit, Top-p sampling
Examples held back from training and used only to measure how well a model does on data it has never seen.
Taught in The learning loopRelated: Training data, Validation set, Overfitting
Computation a model spends while answering, rather than during training, such as writing out reasoning, sampling several answers, or searching over candidates. More of it often improves accuracy on hard problems.
Taught in Thinking longerRelated: Chain of thought, Self-consistency
The unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.
Taught in TokensRelated: Large language model, Logit
The program that splits text into tokens and maps each one to an id in a vocabulary, and turns ids back into text.
Taught in TokensRelated: Token, Vocabulary, Byte-pair encoding
Letting a language model request actions from external programs, such as searches, calculations, or API calls. The model outputs a structured call; the application runs it and returns the result.
Taught in Models that use toolsRelated: Agent, Prompt injection
Sampling the next token from only the k most probable tokens, after setting all others to zero probability and renormalizing.
Taught in Predicting the next tokenRelated: Top-p sampling, Temperature, Greedy decoding
Sampling the next token from the smallest set of most probable tokens whose probabilities add up to at least p. Also called nucleus sampling.
Taught in Predicting the next tokenRelated: Top-k sampling, Temperature
The examples, usually with known answers, that a model learns from. A model can only learn patterns that are present in its training data.
Related: Machine learning, Model
A neural network architecture built from stacked layers of attention and feed-forward networks, introduced by Vaswani et al. in 2017. Most modern language models are transformers.
Taught in Attention, step by stepRelated: Attention, Large language model, Multi-head attention
When a model is too simple to capture the real pattern, so it does poorly on both the training data and new data.
Taught in Overfitting and generalizationRelated: Overfitting, Generalization
Learning from examples that have no answers attached, by finding structure such as groups or directions of variation in the data itself.
Taught in Finding groups: k-meansRelated: Supervised learning, k-means
A second held-back set used while building a model to choose settings such as the learning rate or model size, so the test set stays untouched until the final check.
Taught in Overfitting and generalizationRelated: Test set, Hyperparameter
An ordered list of numbers. It can also be read as an arrow, or a point, in a space with one dimension per number.
Taught in Vectors and similarityRelated: Embedding, Dot product, Cosine similarity
A database built to store embeddings and quickly find the ones most similar to a query vector, usually with an approximate nearest neighbor index.
Taught in Search by meaningRelated: Semantic search, Retrieval-augmented generation
The fixed list of tokens a tokenizer and model know. Each token's position in the list is its id.
Taught in TokensRelated: Token, Tokenizer, Byte-pair encoding
A number that sets how much one input matters to a neuron, and in which direction. Positive weights push the output up, negative weights push it down.
Taught in The artificial neuronRelated: Artificial neuron, Bias (neuron), Parameter
Doing a task without any examples of that specific task during training or in the prompt, for example classifying images into categories the model was never trained to label.
Taught in One space for images and textRelated: Contrastive learning