Compute and scaling laws

Choose a model size and a number of training tokens, see the compute it costs, and compare it with the compute-optimal split under two fitted scaling laws.

AdvancedExplained in How LLMs are trained

Spend a compute budget

Pick a model size and a number of training tokens. The curves show the fitted loss of every other way to spend the same compute.

  • Replication fit (Besiroglu et al. 2024)
  • Original fit as printed (Hoffmann et al. 2022)
  • Your model
Best use of this budget, replication fit
73.0B parameters on 1.34T tokens, about 18 tokens per parameter
Best use of this budget, original fit
32.5B parameters on 3.02T tokens, about 93 tokens per parameter
70.0B
1.40T

Published training runs

Training compute, 6ND5.9 x 1023 FLOP
Tokens per parameter20
Fitted loss, replication fit1.974

Try this

  • Press each published training run and see where it sits relative to the minimum for its own budget.
  • Hover the curves to read off the loss of every other split of the same compute.
  • Notice how far apart the two fits put the best model size, and read the lesson for why.

Try "embedding", "softmax", "overfitting", or "backpropagation".