Label Smoothing is a regularization technique that softens hard one-hot labels by distributing a small amount of probability to non-target classes — instead of training with labels $[0, 0, 1, 0]$, use $[epsilon/K, epsilon/K, 1-epsilon, epsilon/K]$, preventing the model from becoming overconfident.
Label Smoothing Formulation
- Smoothed Label: $y_s = (1 - epsilon) cdot y_{one-hot} + epsilon / K$ where $K$ is the number of classes.
- $epsilon$ Parameter: Typically 0.05-0.1 — small enough to preserve the correct class, large enough to regularize.
- Effect: The model learns to predict ~90% for the correct class instead of trying to reach 100%.
- Calibration: Label smoothing improves model calibration — predicted probabilities better reflect true confidence.
Why It Matters
- Overconfidence: Without smoothing, models become extremely overconfident — label smoothing prevents this.
- Generalization: Acts as a regularizer — improves generalization by preventing the model from fitting hard labels exactly.
- Standard Practice: Used in most modern image classification (ResNet, EfficientNet, ViT) and NLP (BERT, GPT).
Label Smoothing is humble predictions — preventing overconfidence by teaching the model that no class should be predicted with 100% certainty.
label smoothingmachine learning
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.