Lesson 1 of 8

Predicting the next token

At every step a language model scores every token in its vocabulary. Softmax turns the scores into probabilities, and a sampling rule picks one. Temperature, top-k, and top-p shape that choice.

Advanced16 min

In this lesson you will

  • Describe what a language model outputs at each step, one score per token in its vocabulary
  • Turn logits into probabilities with softmax and predict what temperature does to them
  • Compare greedy decoding, top-k, and top-p sampling, and choose between them

A Large language modelA neural network, usually a transformer with billions of parameters, trained on large amounts of text to predict the next token. Chatbots are built on them.Open in glossary writes one TokenThe unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.Open in glossary at a time. Given everything so far, it produces a probability for every possible next token, one token is chosen, it is appended to the text, and the whole process repeats. Generating this way, where each output becomes part of the next input, is called Autoregressive generationProducing a sequence one element at a time, where each new element is predicted from all the ones before it and then appended to the input.Open in glossary generation.

Everything else in this track, attention, position encodings, the transformer block, exists to make that next-token distribution good. This lesson covers the last step: how a distribution becomes an actual word on your screen.

Scores for every token

A model’s final layer does not output a word. It outputs a vector of real numbers called LogitA raw, unnormalized score a model outputs for one option, such as one token in the vocabulary. Softmax turns a vector of logits into probabilities.Open in glossary, one for every token in its vocabulary. GPT-2’s vocabulary has 50,257 tokens; Llama 3’s has 128,256. So at every step the model produces tens of thousands of scores, and higher means “more likely to come next”.

Logits can be any real number, positive or negative. To get probabilities, the model applies SoftmaxA function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.Open in glossary:

pi=ezi/T∑jezj/Tp_i = \frac{e^{z_i / T}}{\sum_{j} e^{z_j / T}}

where ziz_i is the logit for token ii and TT is the TemperatureA setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.Open in glossary, which is 1 unless you change it. The exponential makes every value positive, and dividing by the sum makes them add up to 1.

One property matters a lot in practice: softmax only cares about differences between logits. Adding the same number to every logit changes nothing, and the ratio between two tokens’ probabilities is

pipj=e(zi−zj)/T.\frac{p_i}{p_j} = e^{(z_i - z_j)/T}.

Shape the distribution, then draw

The demo shows twelve candidate next tokens for three prompts. Every control changes the distribution; the buttons draw from it.

From scores to a chosen token

The bars show the probability of each candidate after temperature, top-k, and top-p. Then one token is drawn at random.

The capital of France is?
Candidate next tokens with their probabilities after your settings
TokenProbabilitypDraws
␣Paris77.9%
␣a7.1%
␣the5.2%
␣located2.6%
␣known1.7%
␣home1.4%
␣one1.2%
␣not0.9%
␣also0.7%
␣called0.6%
␣beautiful0.5%
␣Lyon0.3%

after your settings at temperature 1, before any cuts

Prompt
1.00
off
off
Tokens still in play12 of 12
Uncertainty (entropy)1.39 bits
Most likely␣Paris 77.9%

Try this

  • On Fact at temperature 1, press Sample 100 times. “Paris” wins most draws, but not all of them. Before you press it again, guess how many times a token with 1% probability will appear.
  • Drag temperature to 0. Every bar but one drops to zero. This is Greedy decodingGenerating text by always choosing the single most probable next token. Equivalent to sampling at temperature 0.Open in glossary: always take the most likely token.
  • Switch to Story and raise temperature to 1.5. The entropy readout climbs as the bars flatten. The thin ticks show where each bar was at temperature 1.
  • Set top-p to 0.5 on Story and count the survivors. Then switch to Fact. Top-p keeps far fewer tokens when the model is confident; top-k keeps the same number either way.

Notice the space at the start of every candidate. Real tokenizers usually attach the leading space to the word that follows it, so ” Paris” and “Paris” are different tokens. The Language as vectors track covers how text is split into tokens.

What temperature does

Temperature divides every logit before the softmax. Because the probability ratio is e(zi−zj)/Te^{(z_i - z_j)/T}:

  • T<1T < 1 sharpens. Halving the temperature squares every odds ratio. A token that was 2.7 times as likely as another becomes about 7.4 times as likely.
  • T>1T > 1 flattens. Doubling the temperature takes the square root of every odds ratio, so unlikely tokens get more of a chance.
  • T→0T \to 0 puts all probability on the top logit (greedy), and T→∞T \to \infty approaches a uniform distribution.

The order of the tokens never changes. Temperature only decides how strongly the ranking is enforced.

Cutting off the tail: top-k and top-p

Even when each unlikely token has a tiny probability, a vocabulary has tens of thousands of them, and together they can hold a noticeable share. Over hundreds of generated tokens, a draw from that long tail becomes likely, and one strange token can send the rest of the text off course. Truncation methods remove the tail before sampling:

  • Top-k samplingSampling the next token from only the k most probable tokens, after setting all others to zero probability and renormalizing.Open in glossary keeps only the kk most probable tokens, then renormalizes. It was popularized for story generation by Fan et al. (2018). Its weakness is that kk is fixed: 40 tokens may be far too many when the model is sure, and too few when many continuations are reasonable.
  • Top-p samplingSampling the next token from the smallest set of most probable tokens whose probabilities add up to at least p. Also called nucleus sampling.Open in glossary, also called nucleus sampling (Holtzman et al., “The Curious Case of Neural Text Degeneration”, 2019), keeps the smallest set of most probable tokens whose probabilities add up to at least pp, then renormalizes. The set grows and shrinks with the model’s confidence.

Hugging Face Transformers and vLLM apply these in a fixed order: temperature first, then top-k, then top-p, then renormalize and draw. The demo follows the same order. Not every library agrees: llama.cpp applies temperature last by default, which changes which tokens top-p keeps.

MethodWhat it keepsTypical use
Greedy (T=0T = 0)The single most likely tokenShort factual answers, some code; deterministic in principle
Pure samplingEverythingRarely used alone for long text
Top-kThe kk most likely tokensSimple and fast; fixed size regardless of confidence
Top-pThe smallest set reaching mass ppA common default for open-ended text, often with TT near 0.7 to 1
Why logits are log probabilitiesOptional

Language models are trained to maximize the probability they assign to the actual next token in their training text, which is the same as minimizing the cross-entropy loss −log⁡pcorrect-\log p_{\text{correct}}. Taking the log of the softmax gives

log⁡pi=zi−log⁡∑jezj,\log p_i = z_i - \log \sum_j e^{z_j},

so each logit is the log probability of its token, up to a constant shared by all tokens. That is why logits can be any real number, why only their differences matter, and why “log-probability” and “logit” are used almost interchangeably when people inspect models.

Greedy decoding is deterministic on paper. In deployed systems, tiny floating point differences from batching and parallel hardware can occasionally change which of two nearly tied tokens wins, so outputs at temperature 0 are not always bit-for-bit repeatable.

Key ideas

  • A language model outputs one logit per vocabulary token at every step; softmax turns them into probabilities.
  • Only differences between logits matter. The odds between two tokens are e(zi−zj)/Te^{(z_i - z_j)/T}.
  • Temperature sharpens (T<1T < 1) or flattens (T>1T > 1) the distribution without changing the ranking. T=0T = 0 is greedy decoding.
  • Top-k keeps a fixed number of tokens; top-p keeps a fixed amount of probability, so it adapts to the model’s confidence.
  • Generation repeats this choice once per token, which is why one bad draw can affect everything after it.

Check yourself

Pick an answer to see why it is right or wrong. Nothing is graded. Your first answer is saved in this browser so the question can come back for review.

1A model outputs logits [2, 1, 0] for three tokens. You add 5 to every logit. What happens to the probabilities?
2Two tokens have logits that differ by 1. At temperature 1 the more likely one is about 2.7 times as likely as the other. About how many times as likely is it at temperature 0.5?
3With top-p set to 0.9, why does a confident distribution keep fewer tokens than a flat one?

Progress is saved in this browser only.

Up nextAttention, step by step
Next
Transformers and LLMs
  1. 1Predicting the next token
  2. 2Attention, step by step
  3. 3Many heads and the causal mask
  4. 4Where words are
  5. 5The transformer block
  6. 6Inside a real language model
  7. 7How LLMs are trained
  8. 8Generating fast: the KV cache

Try "embedding", "softmax", "overfitting", or "backpropagation".