Spend a compute budget
Pick a model size and a number of training tokens. The curves show the fitted loss of every other way to spend the same compute.
- Replication fit (Besiroglu et al. 2024)
- Original fit as printed (Hoffmann et al. 2022)
- Your model
- Best use of this budget, replication fit
- 73.0B parameters on 1.34T tokens, about 18 tokens per parameter
- Best use of this budget, original fit
- 32.5B parameters on 3.02T tokens, about 93 tokens per parameter
Published training runs
Training compute, 6ND5.9 x 1023 FLOP
Tokens per parameter20
Fitted loss, replication fit1.974
Try this
- Press each published training run and see where it sits relative to the minimum for its own budget.
- Hover the curves to read off the loss of every other split of the same compute.
- Notice how far apart the two fits put the best model size, and read the lesson for why.