label smoothing

**Label Smoothing** is a **regularization technique that softens hard one-hot labels by distributing a small amount of probability to non-target classes** — instead of training with labels $[0, 0, 1, 0]$, use $[epsilon/K, epsilon/K, 1-epsilon, epsilon/K]$, preventing the model from becoming overconfident. **Label Smoothing Formulation** - **Smoothed Label**: $y_s = (1 - epsilon) cdot y_{one-hot} + epsilon / K$ where $K$ is the number of classes. - **$epsilon$ Parameter**: Typically 0.05-0.1 — small enough to preserve the correct class, large enough to regularize. - **Effect**: The model learns to predict ~90% for the correct class instead of trying to reach 100%. - **Calibration**: Label smoothing improves model calibration — predicted probabilities better reflect true confidence. **Why It Matters** - **Overconfidence**: Without smoothing, models become extremely overconfident — label smoothing prevents this. - **Generalization**: Acts as a regularizer — improves generalization by preventing the model from fitting hard labels exactly. - **Standard Practice**: Used in most modern image classification (ResNet, EfficientNet, ViT) and NLP (BERT, GPT). **Label Smoothing** is **humble predictions** — preventing overconfidence by teaching the model that no class should be predicted with 100% certainty.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account