Sparse models activate only a subset of parameters for each input, enabling larger total capacity with fixed compute. Core idea: Route each input to subset of model (experts), rest of parameters inactive. More total parameters without proportional compute increase. Mixture of Experts (MoE): Predominant sparse architecture. Router selects which experts process each token. Sparsity patterns: Expert-based (MoE), unstructured sparsity (zero weights), attention sparsity (attend to subset of tokens). Efficiency gain: 8x7B MoE has 56B total params but activates only 7B per token. Compute of 7B, capacity approaching 56B. Training challenges: Load balancing (experts used equally), routing stability, communication overhead in distributed training. Inference considerations: Need all parameters in memory even if not all active. Different compute vs memory trade-off than dense. Examples: Mixtral 8x7B, GPT-4 (rumored), Switch Transformer, GShard. Advantages: Scale capacity without proportional compute, potential for specialization. Disadvantages: More complex, less predictable, some routing overhead. Increasingly important for frontier models.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.