Random Erasing in ViT is a data augmentation technique that randomly masks rectangular patches in input images during Vision Transformer training to improve robustness and reduce overfitting.
What Is Random Erasing?
- Method: Replace random image regions with random values or mean pixel
- Parameters: Probability, area ratio (0.02-0.4), aspect ratio
- Effect: Forces model to learn from partial information
- Origin: Zhong et al. 2017, widely adopted in ViT training
Why Random Erasing Matters
ViTs can overfit to specific image regions. Random erasing encourages attention to diverse features and improves generalization.
<svg viewBox="0 0 435 245" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto" role="img"><rect x="0" y="0" width="435" height="245" rx="12" fill="#0d1117"/><g font-family="ui-monospace,SFMono-Regular,Menlo,Consolas,"Liberation Mono",monospace" font-size="14"><text xml:space="preserve" x="20" y="31.7"><tspan fill="#c9d1d9">Random Erasing Example:</tspan></text><text xml:space="preserve" x="20" y="50.7"><tspan fill="#c9d1d9">Original Image: After Random Erasing:</tspan></text><text xml:space="preserve" x="20" y="69.7"><tspan fill="#6e7681">┌─────────────────┐</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">┌─────────────────┐</tspan></text><text xml:space="preserve" x="20" y="88.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> 🐱 </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> 🐱 </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="107.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Cat face </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Cat███ </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="126.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">→</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> ███ </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="145.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Body </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> Body </tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="164.7"><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">│</tspan><tspan fill="#c9d1d9"> ███████</tspan><tspan fill="#6e7681">│</tspan></text><text xml:space="preserve" x="20" y="183.7"><tspan fill="#6e7681">└─────────────────┘</tspan><tspan fill="#c9d1d9"> </tspan><tspan fill="#6e7681">└─────────────────┘</tspan></text><text xml:space="preserve" x="20" y="202.7"></text><text xml:space="preserve" x="20" y="221.7"><tspan fill="#c9d1d9">Model must recognize cat without erased patches</tspan></text></g></svg>
Random Erasing in ViT Recipe:
transforms.Compose([
transforms.RandomResizedCrop(224),
transforms.RandomHorizontalFlip(),
transforms.ToTensor(),
RandomErasing(
probability=0.25,
sl=0.02, sh=0.4, # area ratio
r1=0.3, # aspect ratio min
),
])
Typical improvement: +0.5-1.5% top-1 accuracy on ImageNet.
random erasingvit augmentationimage masking
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.