A dense language model uses every one of its ParameterA number inside a model that training adjusts, such as the slope of a line or a weight in a neural network. Large language models have billions of them.Open in glossary for every token it reads or writes. That gets expensive quickly. A Mixture of expertsAn architecture with many alternative sub-networks (experts) and a router that runs only a few of them for each token, so a model can store many parameters while using few per token.Open in glossary (MoE) model takes a different approach: it holds many alternative sub-networks, called experts, and for each token it runs only a few of them. The model can store far more knowledge than it spends computation on.
The idea is old (Jacobs and colleagues described mixtures of local experts in 1991) and was scaled up for deep networks by Shazeer et al. (2017) with a sparsely gated layer. Many recent large models are built this way, including Mixtral 8x7B and DeepSeek-V3.
Where the experts live
In a transformer, each layer has an attention block followed by a feed-forward network. In an MoE transformer the attention block stays as it is, and the single feed-forward network is replaced by several experts, each a feed-forward network of its own, plus a small RouterThe small learned layer in a mixture-of-experts model that scores every expert for each token and picks which few experts will process it.Open in glossary.
For a token with hidden vector , the router computes one score per expert, . The top scores pick the experts that will run. Their outputs are mixed with gate weights that are a softmax over just the chosen scores:
This is how Mixtral combines its experts. Experts outside the top contribute nothing and cost nothing for that token.
A router picks experts for each token
Shading is the router's probability for each expert. Outlined cells are the experts that actually run.
Try this
- Click through the tokens. Different tokens light up different pairs of experts, and the formula shows how their outputs are weighted.
- Switch to Top 1. Now each token gets exactly one expert with weight 1, as in the Switch Transformer (Fedus et al., covered below).
- Watch the load bars. Some experts get several tokens and some get none. In a real model, an unused expert is wasted memory, and an overloaded one is a bottleneck.
- Compare 4 experts with 8 experts at top 2. The work per token is the same, but the number of possible expert pairs grows from 6 to 28.
Keeping the load balanced
Left alone, a router tends to favor a few experts. Those experts improve faster because they see more tokens, so the router favors them even more, and the rest wither. MoE training adds an auxiliary loss to push against this. A common form, from the Switch Transformer (Fedus et al., 2022), is
where is the number of experts, is the fraction of token assignments that went to expert , is the router’s average probability for expert , and is a small coefficient. Without , the sum is 1 when routing is perfectly even and grows as tokens crowd onto a few experts. The demo shows it as the balance term. Systems also cap how many tokens one expert can take per batch, and newer designs such as DeepSeek-V3 use other balancing schemes.
Total parameters versus active parameters
Because only of experts run, an MoE model has two sizes. The total count is everything stored. The active count is what one token actually passes through: all the shared parts plus experts per layer.
Total versus active parameters
Same shape as Mixtral 8x7B: 32 layers, model width 4,096, expert hidden size 14,336. Change the experts.
Shared by every token: 1.6B Expert weights: 0.2B per expert per layer x 32 layers
Try this
- At the default settings (8 experts, top 2) the counts match Mixtral 8x7B’s published figures, about 46.7B total and 12.9B active.
- Set experts to 1 and k to 1. That is a dense model of about 7B parameters, which is where the “7B” in the name comes from.
- Raise the experts to 64 with k = 2. Total parameters climb past 300B while active parameters barely move (only the small router grows).
The name “8x7B” suggests 56B parameters, but only the feed-forward blocks are copied eight times. Attention, embeddings, and normalization are shared, so the total is about 47B. For each token, the model does roughly the arithmetic of a 13B dense model.
Counting the parameters yourselfOptional
For one layer with model width , expert hidden size , attention heads, and key-value heads:
- Each gated (SwiGLU) expert has three weight matrices: parameters.
- Attention has query and output projections of each, plus key and value projections of each, for grouped-query attention.
- The router is a matrix.
Mixtral 8x7B uses , , 32 layers, 32 heads with 8 key-value heads, and a 32,000-token vocabulary with separate input and output embeddings. One expert is parameters, so the experts alone are . Attention adds about 1.34B, embeddings about 0.26B, and the router and norms about 1.3 million, giving about 46.7B. With two experts per token the expert share drops to about 11.3B, for about 12.9B active.
What mixture of experts buys, and what it costs
Compute. Training and inference arithmetic per token scale with active parameters, so an MoE model can be trained on more data for the same budget, or be larger for the same speed.
Memory. Every expert must be stored, and typically held in accelerator memory, because the next token might need any of them. Mixtral’s 46.7B parameters take about 93 GB at 16 bits per weight, even though each token uses only 12.9B of them.
Engineering. Tokens for different experts must be sent to wherever those experts live, often across many chips. Load imbalance turns directly into idle hardware.
Key ideas
- An MoE layer replaces one feed-forward network with many experts and a router that picks the top for each token.
- The outputs of the chosen experts are mixed with gate weights that sum to 1.
- Total parameters count everything stored; active parameters count what one token uses. Compute follows active, memory follows total.
- Training adds a balancing pressure so that tokens spread across experts.
- “Experts” are architectural slots; in trained models they do not line up neatly with human topics.