Mixture of Experts (MoE) is the neural network architecture that routes each input token to a subset of specialized "expert" sub-networks through a learned gating function — enabling models with trillions of parameters while only activating a fraction of them per forward pass, achieving the capacity of dense models at a fraction of the compute cost and making efficient scaling beyond dense model limits practical.
Core Architecture
A standard MoE layer replaces the dense feed-forward network (FFN) in a Transformer block with N parallel expert FFNs and a gating (router) network:
- Experts: N independent FFN sub-networks (typically 8-128), each with identical architecture but separate learned weights.
- Router/Gate: A small network (usually a linear layer + softmax) that takes the input token and produces a probability distribution over experts. The top-K experts (typically K=1 or K=2) are selected for each token.
- Sparse Activation: Only the selected K experts process each token. Total model parameters scale with N (number of experts), but compute per token scales with K — independent of N.
Gating Mechanisms
- Top-K Routing: Select the K experts with highest gate probability. Multiply each expert's output by its gate weight and sum. Simple and effective but prone to load imbalance (popular experts get most tokens).
- Switch Routing: K=1 (single expert per token). Maximum sparsity and simplest implementation. Used in Switch Transformer (Google, 2021) achieving 7x training speedup over T5-Base at equivalent FLOPS.
- Expert Choice Routing: Instead of tokens choosing experts, each expert selects its top-K tokens. Guarantees perfect load balance but changes the computation graph (variable tokens per sequence position).
Load Balancing
The critical engineering challenge. Without intervention, a few experts receive most tokens (rich-get-richer collapse), wasting the capacity of idle experts:
- Auxiliary Loss: Add a loss term penalizing uneven expert utilization. The standard approach — a small coefficient (0.01-0.1) balances routing diversity against task performance.
- Expert Capacity Factor: Each expert processes at most C × (N_tokens / N_experts) tokens per batch. Tokens exceeding capacity are dropped or rerouted.
- Random Routing: Mix deterministic top-K selection with random assignment to ensure exploration of all experts during training.
Scaling Results
- GShard (Google, 2020): 600B parameter MoE with 2048 experts across 2048 TPU cores.
- Switch Transformer (2021): Demonstrated scaling to 1.6T parameters with simple top-1 routing.
- Mixtral 8x7B (Mistral, 2023): 8 experts, 2 active per token. 47B total parameters, 13B active — matching or exceeding LLaMA-2 70B quality at 6x lower inference cost.
- DeepSeek-V3 (2024): 671B total parameters, 37B active per token. MoE enabling frontier-quality at dramatically reduced training cost.
Inference Challenges
MoE models require all expert weights in memory (or fast-swappable) even though only K are active per token. For Mixtral 8x7B: 47B parameters in memory for 13B-equivalent compute. Expert parallelism distributes experts across GPUs, but routing decisions create all-to-all communication patterns that stress interconnect bandwidth.
Mixture of Experts is the architectural paradigm that breaks the linear relationship between model quality and inference cost — proving that scaling model capacity through conditional computation produces better results per FLOP than scaling dense models, and enabling the next generation of frontier language models.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.