A Large language modelA neural network, usually a transformer with billions of parameters, trained on large amounts of text to predict the next token. Chatbots are built on them.Open in glossary writes one TokenThe unit of text a language model reads and writes. A token can be a whole word, part of a word, a single character, or punctuation, often including a leading space.Open in glossary at a time. Given everything so far, it produces a probability for every possible next token, one token is chosen, it is appended to the text, and the whole process repeats. Generating this way, where each output becomes part of the next input, is called Autoregressive generationProducing a sequence one element at a time, where each new element is predicted from all the ones before it and then appended to the input.Open in glossary generation.
Everything else in this track, attention, position encodings, the transformer block, exists to make that next-token distribution good. This lesson covers the last step: how a distribution becomes an actual word on your screen.
Scores for every token
A model’s final layer does not output a word. It outputs a vector of real numbers called LogitA raw, unnormalized score a model outputs for one option, such as one token in the vocabulary. Softmax turns a vector of logits into probabilities.Open in glossary, one for every token in its vocabulary. GPT-2’s vocabulary has 50,257 tokens; Llama 3’s has 128,256. So at every step the model produces tens of thousands of scores, and higher means “more likely to come next”.
Logits can be any real number, positive or negative. To get probabilities, the model applies SoftmaxA function that turns a list of real-valued scores into probabilities that are positive and sum to 1, by exponentiating each score and dividing by the total.Open in glossary:
where is the logit for token and is the TemperatureA setting for how adventurous a model's choices are. Low temperature sticks to the likeliest options; high temperature spreads choices toward less likely ones. Technically, it divides the logits before softmax.Open in glossary, which is 1 unless you change it. The exponential makes every value positive, and dividing by the sum makes them add up to 1.
One property matters a lot in practice: softmax only cares about differences between logits. Adding the same number to every logit changes nothing, and the ratio between two tokens’ probabilities is
Shape the distribution, then draw
The demo shows twelve candidate next tokens for three prompts. Every control changes the distribution; the buttons draw from it.
From scores to a chosen token
The bars show the probability of each candidate after temperature, top-k, and top-p. Then one token is drawn at random.
| Token | Probability | p | Draws |
|---|---|---|---|
| ␣Paris | 77.9% | ||
| ␣a | 7.1% | ||
| ␣the | 5.2% | ||
| ␣located | 2.6% | ||
| ␣known | 1.7% | ||
| ␣home | 1.4% | ||
| ␣one | 1.2% | ||
| ␣not | 0.9% | ||
| ␣also | 0.7% | ||
| ␣called | 0.6% | ||
| ␣beautiful | 0.5% | ||
| ␣Lyon | 0.3% |
after your settings at temperature 1, before any cuts
Try this
- On Fact at temperature 1, press Sample 100 times. “Paris” wins most draws, but not all of them. Before you press it again, guess how many times a token with 1% probability will appear.
- Drag temperature to 0. Every bar but one drops to zero. This is Greedy decodingGenerating text by always choosing the single most probable next token. Equivalent to sampling at temperature 0.Open in glossary: always take the most likely token.
- Switch to Story and raise temperature to 1.5. The entropy readout climbs as the bars flatten. The thin ticks show where each bar was at temperature 1.
- Set top-p to 0.5 on Story and count the survivors. Then switch to Fact. Top-p keeps far fewer tokens when the model is confident; top-k keeps the same number either way.
Notice the space at the start of every candidate. Real tokenizers usually attach the leading space to the word that follows it, so ” Paris” and “Paris” are different tokens. The Language as vectors track covers how text is split into tokens.
What temperature does
Temperature divides every logit before the softmax. Because the probability ratio is :
- sharpens. Halving the temperature squares every odds ratio. A token that was 2.7 times as likely as another becomes about 7.4 times as likely.
- flattens. Doubling the temperature takes the square root of every odds ratio, so unlikely tokens get more of a chance.
- puts all probability on the top logit (greedy), and approaches a uniform distribution.
The order of the tokens never changes. Temperature only decides how strongly the ranking is enforced.
Cutting off the tail: top-k and top-p
Even when each unlikely token has a tiny probability, a vocabulary has tens of thousands of them, and together they can hold a noticeable share. Over hundreds of generated tokens, a draw from that long tail becomes likely, and one strange token can send the rest of the text off course. Truncation methods remove the tail before sampling:
- Top-k samplingSampling the next token from only the k most probable tokens, after setting all others to zero probability and renormalizing.Open in glossary keeps only the most probable tokens, then renormalizes. It was popularized for story generation by Fan et al. (2018). Its weakness is that is fixed: 40 tokens may be far too many when the model is sure, and too few when many continuations are reasonable.
- Top-p samplingSampling the next token from the smallest set of most probable tokens whose probabilities add up to at least p. Also called nucleus sampling.Open in glossary, also called nucleus sampling (Holtzman et al., “The Curious Case of Neural Text Degeneration”, 2019), keeps the smallest set of most probable tokens whose probabilities add up to at least , then renormalizes. The set grows and shrinks with the model’s confidence.
Hugging Face Transformers and vLLM apply these in a fixed order: temperature first, then top-k, then top-p, then renormalize and draw. The demo follows the same order. Not every library agrees: llama.cpp applies temperature last by default, which changes which tokens top-p keeps.
| Method | What it keeps | Typical use |
|---|---|---|
| Greedy () | The single most likely token | Short factual answers, some code; deterministic in principle |
| Pure sampling | Everything | Rarely used alone for long text |
| Top-k | The most likely tokens | Simple and fast; fixed size regardless of confidence |
| Top-p | The smallest set reaching mass | A common default for open-ended text, often with near 0.7 to 1 |
Why logits are log probabilitiesOptional
Language models are trained to maximize the probability they assign to the actual next token in their training text, which is the same as minimizing the cross-entropy loss . Taking the log of the softmax gives
so each logit is the log probability of its token, up to a constant shared by all tokens. That is why logits can be any real number, why only their differences matter, and why “log-probability” and “logit” are used almost interchangeably when people inspect models.
Greedy decoding is deterministic on paper. In deployed systems, tiny floating point differences from batching and parallel hardware can occasionally change which of two nearly tied tokens wins, so outputs at temperature 0 are not always bit-for-bit repeatable.
Key ideas
- A language model outputs one logit per vocabulary token at every step; softmax turns them into probabilities.
- Only differences between logits matter. The odds between two tokens are .
- Temperature sharpens () or flattens () the distribution without changing the ranking. is greedy decoding.
- Top-k keeps a fixed number of tokens; top-p keeps a fixed amount of probability, so it adapts to the model’s confidence.
- Generation repeats this choice once per token, which is why one bad draw can affect everything after it.