Gradient accumulation simulates larger batch sizes by accumulating gradients over multiple forward-backward passes before updating. How it works: Run forward and backward multiple times, sum gradients, then apply single optimizer step. Effective batch = micro-batch x accumulation steps. Why useful: GPU memory limits batch size. Want larger effective batch for training stability without more memory. Implementation: Call loss.backward() multiple times, then optimizer.step() and zero_grad(). Or use framework support. Memory benefit: Same memory as small batch, but large batch training dynamics. Training dynamics: Large batches often need learning rate scaling (linear scaling rule). May affect convergence. Trade-off: More forward/backward passes before update = slower wall-clock time. Worthwhile when batch size matters. Common use cases: Limited GPU memory, matching batch size across different hardware, very large batch training experiments. Distributed training: Accumulation within device, sync gradients after accumulation steps. Reduces communication frequency. Best practices: Scale learning rate appropriately, consider gradient normalization, validate against true large batch training.
Related Topics
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.