Mixture of Experts (MoE) is the neural network architecture that routes each input token through only a subset of specialized sub-networks (experts) selected by a learned gating mechanism — enabling models with trillions of parameters while keeping per-token computation constant, because only 1-2 experts out of hundreds are activated for any given input.
The Scaling Dilemma MoE Solves
Dense transformer models scale by increasing width (hidden dimension) and depth (layers), but compute cost grows proportionally with parameter count. A 1.8T parameter dense model would require enormous FLOPs per token. MoE decouples parameter count from compute cost: a 1.8T MoE model with 128 experts and top-2 routing activates only ~28B parameters per token — the same compute as a 28B dense model but with access to a much larger knowledge capacity.
Architecture
In a typical MoE transformer, every other feed-forward network (FFN) layer is replaced with an MoE layer:
- Experts: N identical FFN sub-networks (e.g., N=8, 64, or 128), each with independent parameters.
- Router (Gating Network): A lightweight linear layer that takes the token representation as input and outputs a probability distribution over experts. The top-K experts (typically K=1 or K=2) are selected per token.
- Combination: The outputs of the selected experts are weighted by their gating probabilities and summed.
Load Balancing Challenge
Without constraints, the router tends to collapse — sending all tokens to a few popular experts while others remain unused. This wastes capacity and creates compute imbalance across devices (each expert is placed on a different GPU). Solutions:
- Auxiliary Load Balancing Loss: An additional loss term that penalizes uneven expert utilization, encouraging the router to distribute tokens evenly.
- Expert Capacity Factor: Each expert has a maximum number of tokens it can process per batch. Overflow tokens are either dropped or routed to a shared fallback expert.
- Token Choice vs. Expert Choice: In expert-choice routing, each expert selects its top-K tokens rather than each token selecting its top-K experts — guaranteeing perfect load balance.
Training Infrastructure
MoE layers require expert parallelism: experts are distributed across GPUs, and all-to-all communication shuffles tokens to their assigned expert's GPU and back. This all-to-all pattern is bandwidth-intensive and requires careful overlap with computation. Frameworks like Megatron-LM and DeepSpeed-MoE provide optimized implementations combining data, tensor, expert, and pipeline parallelism.
Notable MoE Models
- Switch Transformer (Google): Top-1 routing with simplified load balancing. Demonstrated 7x training speedup over dense T5 at equivalent compute.
- Mixtral 8x7B (Mistral): 8 experts per layer, top-2 routing. 46.7B total parameters but ~13B active per token. Outperforms LLaMA 2 70B at much lower inference cost.
- DeepSeek-V2/V3: MoE with fine-grained experts (up to 256) and shared expert layers for common knowledge.
Mixture of Experts is the architectural paradigm that breaks the linear relationship between model capacity and inference cost — enabling foundation models to store vastly more knowledge in their parameters while maintaining practical serving latency and throughput.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.