Home Knowledge Base Motivation

Model parallelism splits model layers across devices, enabling training models too large for single GPU memory. Motivation: Models like GPT-3 (175B) exceed single GPU memory. Must distribute parameters across devices. Tensor parallelism: Split individual layers across devices. Matrix multiplications distributed, results combined. Megatron-LM style. Layer parallelism: Different layers on different devices. Simpler but less communication overlap. How tensor parallelism works: For linear layer Y = XW, split W column-wise across devices. Each computes partial result, all-reduce to combine. Communication overhead: Requires synchronization within layers. Latency-sensitive, works best with fast interconnects (NVLink). Memory benefit: Each device stores fraction of parameters. 8-way tensor parallel = 1/8 memory per device. Trade-offs: More communication than data parallelism, efficiency depends on interconnect speed, implementation complexity. When to use: Model doesnt fit on single device, have fast interconnects, need memory distribution. Common setup: Tensor parallel within node (NVLink), data parallel across nodes (Ethernet/InfiniBand). Frameworks: Megatron-LM, DeepSpeed, FairScale, NeMo.

model parallelismmodel training

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.