Home Knowledge Base Masked image modeling (MIM)

Masked image modeling (MIM) is the self-supervised training paradigm where a model reconstructs hidden image patches from visible context - this forces ViT encoders to learn semantic and structural representations instead of memorizing local texture shortcuts.

What Is Masked Image Modeling?

Why MIM Matters

MIM Variants

Pixel Reconstruction:

Token Reconstruction:

Feature Reconstruction:

Training Flow

Step 1:

Step 2:

Masked image modeling is a versatile and scalable self-supervised framework that teaches ViTs to infer missing visual context from surrounding evidence - it is now a core building block for modern vision pretraining.

masked image modelingmimcomputer vision

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.