Home Knowledge Base Strided Attention

Strided Attention is a sparse attention pattern where each token attends to every $s$-th token in the sequence — creating a dilated attention pattern that efficiently captures long-range dependencies without computing full $O(N^2)$ attention.

How Does Strided Attention Work?

Why It Matters

Strided Attention is dilated convolution for attention — skipping tokens at regular intervals to efficiently reach across the entire sequence.

strided attentionsparse attention

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.