The previous lessons took a transformer apart. Every one of its parameters starts as a random number. This lesson is about how those numbers become a model that writes fluent text, and then one that answers questions helpfully, and about the economics that decide how big the model should be and how much it should read.
Three stages
Modern assistants are trained in stages. Each uses a different kind of data and changes something different.
1Pretraining
- Data
- Trillions of tokens of web pages, books, code, and other text.
- Objective
- Predict the next token at every position.
- Result
- A base model that continues any text it is given.
2Supervised fine-tuning
- Data
- Thousands to millions of example prompts with good responses, written or curated by people.
- Objective
- The same next-token loss, applied to the responses.
- Result
- A model that answers in the format of the examples.
3Preference tuning
- Data
- Pairs of responses to the same prompt, with a judgment of which is better.
- Objective
- Make preferred responses more likely: RLHF with a reward model, or DPO directly.
- Result
- An assistant whose answers people rate as more helpful.
PretrainingThe first and most expensive stage of training a language model, in which it learns to predict the next token across a very large amount of text. The result is a base model.Open in glossary is the expensive part. The model reads text and, at every position, is scored on how much probability it gave the token that actually came next. The loss is the average Cross-entropyA loss for predicted probabilities: minus the log of the probability the model gave to the right answer. Confident wrong answers are punished very heavily.Open in glossary over a sequence of tokens:
No human labels are needed: the text supplies its own answers, and every position in every document is a training example. Repeated over trillions of tokens, this one objective forces the model to pick up grammar, facts, styles of reasoning, and the structure of code, because all of them help predict what comes next. The result is a Base modelA language model after pretraining only. It continues text in whatever style the input suggests, rather than following instructions or answering like an assistant.Open in glossary. Ask it a question and it may continue with three more questions, because that is a plausible continuation of a list of questions. You saw this firsthand in the previous lesson.
Fine-tuningContinuing to train a pretrained model on a smaller, targeted dataset, such as example conversations, to change its behavior. Supervised fine-tuning uses examples of good responses.Open in glossary continues training on a much smaller set of example conversations: a prompt and a good response, written or curated by people. The loss is the same cross-entropy, usually applied only to the response tokens. The model learns the format of an answer.
Preference tuning teaches the model which of several acceptable answers people prefer. It is described in its own section below.
Training compute: 6ND
Training cost is measured in floating point operations (FLOP). For a model with parameters trained on tokens, a good estimate is
The forward pass uses each parameter in about one multiply and one add per token: operations. The backward pass, which computes gradients with respect to both the activations and the weights, costs about twice the forward pass: . Attention over the context adds a little more, usually a small share for large models, and the estimate ignores it (Kaplan et al., 2020). GPT-3, with 174.6 billion parameters (175 billion in its name) trained on 300 billion tokens, comes out at FLOP, the figure its authors reported.
Scaling laws
Kaplan et al. (2020) found that a language model’s loss falls smoothly and predictably as you increase parameters, data, or compute, following power laws over several orders of magnitude. Relationships like these are called Scaling lawsFitted curves that predict how a model's loss falls as parameters, training data, or compute grow. They are empirical, and their constants depend on the architecture, data, and training recipe.Open in glossary. That made it possible to plan: train small models, fit a curve, and predict how a model a thousand times larger would do.
Hoffmann et al. (2022) asked the planning question directly. With a fixed compute budget , should you train a bigger model on less data, or a smaller model on more? They trained over 400 models and fitted
where is the fit’s estimate of the lowest loss any model could reach on that data, the second term is the penalty for having too few parameters, and the third is the penalty for seeing too few tokens. Fixing ties to , so the loss becomes a curve in alone with a single best point. Their headline result was that parameters and tokens should grow together, about 20 training tokens per parameter. To test it they trained Chinchilla, 70 billion parameters on 1.4 trillion tokens, with about the same compute as their earlier 280-billion-parameter Gopher, which had seen only 300 billion tokens. Chinchilla outperformed Gopher across a wide range of benchmarks.
Spend a compute budget
Pick a model size and a number of training tokens. The curves show the fitted loss of every other way to spend the same compute.
- Replication fit (Besiroglu et al. 2024)
- Original fit as printed (Hoffmann et al. 2022)
- Your model
- Best use of this budget, replication fit
- 73.0B parameters on 1.34T tokens, about 18 tokens per parameter
- Best use of this budget, original fit
- 32.5B parameters on 3.02T tokens, about 93 tokens per parameter
Published training runs
Try this
- Press GPT-3. Its marker sits far to the right of both minima: for its budget, the fits say a model a few times smaller, trained on several times more tokens, would have reached lower loss.
- Press Chinchilla. Your marker lands next to the minimum of the replication fit.
- Raise the token count while keeping parameters fixed. The fitted loss falls (watch the readout and the axis labels) and both minima shift right: a bigger budget wants both more parameters and more data.
- Compare the two “best use of this budget” lines underneath. For the same compute they disagree about the ideal model size by roughly a factor of two.
The two curves disagree, and the disagreement is instructive. Hoffmann et al. used three different methods. Two of them pointed to about 20 tokens per parameter, which is how Chinchilla was sized. The third, the fitted formula above with the constants printed in the paper, implies several times more tokens per parameter. (Rounding exaggerates this: the unrounded constants give about 60 tokens per parameter rather than 90, still well above 20.) Besiroglu et al. (2024) reconstructed the data from the paper’s figures, refitted the formula, and found constants consistent with the first two methods, about 20 tokens per parameter. The replication fit is plotted as the solid line.
Deriving the compute-optimal splitOptional
Substitute into the fit and minimize over . Setting the derivative to zero,
gives
When , both exponents are close to one half: double the compute and you should multiply both the parameters and the tokens by about . That is the sense in which they “grow together”. The lab computes with this formula, and this site’s tests check it against a direct numerical search along the curve.
The constants used here are, for Hoffmann et al. (2022), , , , , as printed in the paper; and for Besiroglu et al. (2024), , , , , .
Beyond compute-optimal
Chinchilla’s ratio answers one question: what is the lowest loss for a fixed training budget? It says nothing about the cost of using the model. A model that will answer millions of requests is run far more often than it is trained, and every one of those runs costs about operations per token. A smaller model trained on many more tokens can match a larger compute-optimal model’s quality while being much cheaper to serve.
That is why recent open models sit far to the right of 20:
Training tokens per parameter in published models
- GopherRae et al. (2021)1.1
- GPT-3Brown et al. (2020)1.7
- ChinchillaHoffmann et al. (2022)20
- Llama 3 8Babout 15T tokens, Llama Team (2024)1,868
- SmolLM2-135M2T tokens, model card14,870
Preference tuning
Supervised fine-tuning needs someone to write the ideal answer. It is often easier to compare two answers than to write a perfect one. Preference tuning learns from those comparisons.
RLHFReinforcement learning from human feedback. People compare pairs of model responses, a reward model learns to predict their preferences, and the language model is then trained with reinforcement learning to score well on it.Open in glossary (Christiano et al., 2017), as used for InstructGPT (Ouyang et al., 2022), has two steps after fine-tuning:
- Train a reward model. People compare several responses to the same prompt and rank them (InstructGPT’s labelers ranked 4 to 9 at a time); each ranking yields many better-versus-worse pairs. A separate network is trained so that the preferred response scores higher than the rejected one , using the Bradley-Terry model .
- Optimize against it. The language model generates responses, the reward model scores them, and a reinforcement learning algorithm (PPO, in InstructGPT) updates the language model to raise the score, with a penalty for drifting too far from the fine-tuned model so it does not find strange outputs that fool the reward model.
The result was striking: people preferred outputs from a 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3, despite it having over 100 times fewer parameters.
Direct preference optimizationA way to train a language model on pairs of preferred and rejected responses with a single classification-style loss, without a separate reward model or reinforcement learning.Open in glossary (Rafailov et al., 2023) reaches the same goal without the reward model or the reinforcement learning loop. It shows that the RLHF objective can be optimized with a single loss on preference pairs:
where is the model being trained, is the frozen starting model, and controls how far it may move. The loss raises the probability of the preferred response relative to the rejected one, measured against the reference. It is a classification loss, trained like any other, which made preference tuning much simpler to run.
Key ideas
- Pretraining predicts the next token over trillions of tokens and produces a base model. Fine-tuning and preference tuning shape its behavior with far less data.
- Training compute is about FLOP: 2 per parameter per token forward, 4 backward.
- Scaling laws are fitted curves. For a fixed training budget the Chinchilla analysis puts the best model at about 20 tokens per parameter.
- Models that will be run a lot are deliberately trained far past that ratio, trading training compute for cheaper inference.
- RLHF learns a reward model from comparisons and optimizes against it. DPO optimizes the same kind of objective directly with a classification loss.