Mixture of experts

See a router send each token to its top experts, then compare total and active parameter counts for a Mixtral-shaped model.

AdvancedExplained in Mixture of experts

A router picks experts for each token

Shading is the router's probability for each expert. Outlined cells are the experts that actually run.

Experts in the layer
Experts used per token
Expert computations per token2 of 8
Balance term (1.0 is perfectly even)1.37

Total versus active parameters

Same shape as Mixtral 8x7B: 32 layers, model width 4,096, expert hidden size 14,336. Change the experts.

8
2
Total parameters46.7B
Active per token12.9B
Weights in memory at 16 bits93 GB

Try this

  • Switch between top 1 and top 2 routing and watch the gate weights for a token.
  • In the calculator, set experts and k to 1 to recover the dense model, then add experts and watch only the stored total grow.

Try "embedding", "softmax", "overfitting", or "backpropagation".