Home Knowledge Base Register Tokens

Register Tokens are deliberately inserted, learnable blank placeholder tokens injected directly into the input sequence of a Vision Transformer (ViT) specifically engineered to serve as dedicated mathematical scratchpads that absorb and quarantine toxic "outlier" attention artifacts that would otherwise catastrophically corrupt the meaningful feature representations of the actual image patches.

The Artifact Problem

The Register Solution

Why Registers are Necessary

Standard ViT architectures (DINOv2, ViT-L) exhibit severe attention artifacts once scaled to very large parameter counts and high-resolution inputs. The register mechanism eliminates these artifacts without modifying the fundamental Transformer architecture, yielding substantially cleaner attention maps and measurably improved performance on dense tasks like semantic segmentation and object detection.

Register Tokens are the attention junk drawer — purpose-built mathematical wastebaskets that intercept and quarantine toxic information overflow, ensuring the Transformer's critical attention highways remain clean and focused on the actual visual content.

register tokenscomputer vision

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.