loss spikes

**Loss Spikes** are **sudden, sharp increases in training loss that temporarily disrupt the training process** — the loss dramatically increases for a few steps or epochs, then rapidly recovers, often to a value lower than before the spike, suggesting the model is transitioning between different solution basins. **Loss Spike Characteristics** - **Magnitude**: Can be 2-100× the pre-spike loss — sometimes dramatic increases. - **Recovery**: Loss typically recovers within a few hundred to a few thousand steps. - **Causes**: Large learning rates, numerical instability (fp16 overflow), batch composition, data quality issues, or representation reorganization. - **Beneficial**: Some loss spikes precede improved performance — the model "jumps" to a better region of the loss landscape. **Why It Matters** - **Training Stability**: Loss spikes can derail training if severe — require monitoring and mitigation (gradient clipping, loss scaling). - **LLM Training**: Large language model training frequently experiences loss spikes — especially at scale. - **Learning Signal**: Some spikes indicate the model is learning new, qualitatively different representations — a positive sign. **Loss Spikes** are **turbulence in training** — sudden loss increases that can signal either instability issues or beneficial representation transitions.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account