Home Knowledge Base Dilated Attention

Dilated Attention is the sparse attention pattern that spreads tokens apart to cover a larger effective field while keeping the number of attended neighbors low — by stepping through spatial positions with a stride greater than one, the model reaches distant tokens with fewer dot products, letting it process ultra-high-resolution inputs without computing full dense matrices.

What Is Dilated Attention?

Why Dilated Attention Matters

Dilation Strategies

Constant Dilation:

Progressive Dilation:

Hybrid Dilation:

How It Works / Technical Details

Step 1: Build attention neighborhoods by striding across the flattened spatial grid with step d, gathering keys and values only at those positions through strided slicing.

Step 2: Compute scaled dot product attention on the collected subsets, apply softmax, and gather weighted sums; combine with residual projections and feed-forward blocks.

Comparison / Alternatives

AspectDilated AttentionLocal Window (Swin)Global Attention
CoverageSparse globalLocal with shiftsDense global
ComputeO(N/k^2) depending on dO(Nw^2)O(N^2)
DetailVaries with dFixed by wFull detail
Best UseVery high resolutionBalanced performanceModerate sizes

Tools & Platforms

Dilated attention is the sparse yet powerful lens that makes Vision Transformers see far without lifting a heavy quadratic burden — it jumps over nearby redundancy and attends to tokens that truly matter at distant locations.

dilated attention in visioncomputer vision

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.