A router picks experts for each token
Shading is the router's probability for each expert. Outlined cells are the experts that actually run.
Expert computations per token2 of 8
Balance term (1.0 is perfectly even)1.37
Total versus active parameters
Same shape as Mixtral 8x7B: 32 layers, model width 4,096, expert hidden size 14,336. Change the experts.
Stored (all experts)46.7B
Used per token (k experts)12.9B
Shared by every token: 1.6B Expert weights: 0.2B per expert per layer x 32 layers
Total parameters46.7B
Active per token12.9B
Weights in memory at 16 bits93 GB
Try this
- Switch between top 1 and top 2 routing and watch the gate weights for a token.
- In the calculator, set experts and k to 1 to recover the dense model, then add experts and watch only the stored total grow.