Home Knowledge Base Router network

Expert routing determines which experts process each token in Mixture of Experts architectures. Router network: Small network (often single linear layer) that takes token embedding as input, outputs score for each expert. Routing strategies: Top-k: Select k highest-scoring experts. Common: top-1 (single expert) or top-2 (two experts, combine outputs). Token choice: Each token chooses its experts. Expert choice: Each expert chooses its tokens (better load balance). Soft routing: Weight contributions from all experts by router probabilities. More compute but smoother. Routing decisions: Learned during training. Router learns to specialize experts for different input types. Aux losses: Auxiliary loss terms encourage load balancing, prevent expert collapse. Capacity constraints: Limit tokens per expert to ensure balanced workload. Overflow handling varies. Emergent specialization: Experts often specialize (e.g., punctuation expert, code expert) though not always interpretable. Routing overhead: Router computation is small fraction of total. Main overhead is communication in distributed setting. Research areas: Stable routing, better load balancing, interpretable expert roles.

expert routingmodel architecture

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.