Landmark Attention is the efficient transformer attention mechanism that reduces computational complexity by routing all token attention through a sparse set of landmark (anchor) tokens that serve as information hubs — achieving sub-quadratic attention cost while preserving global information flow — the architecture that demonstrates how strategically placed landmark tokens can serve as a compressed global context, enabling long-sequence processing without the full O(n²) cost of standard self-attention.
What Is Landmark Attention?
- Definition: A modified attention mechanism where regular tokens attend only to nearby local tokens and to a set of specially designated landmark tokens, while landmark tokens attend to all other landmarks — creating a two-level attention hierarchy with O(n × k) complexity where k << n is the number of landmarks.
- Landmark Selection: Landmarks are chosen at fixed intervals (every m-th token), at content boundaries (sentence/paragraph breaks), or through learned prominence scoring — they serve as representative summaries of their local region.
- Two-Level Attention: (1) Local tokens attend to their neighborhood + all landmarks (sparse), (2) Landmarks attend to all other landmarks (dense but small) — global information propagates through the landmark network while local processing remains efficient.
- Information Bridge: Landmarks act as bridges between distant sequence regions — a token at position 1 can influence a token at position 10,000 through their respective nearest landmarks, which are connected via landmark-to-landmark attention.
Why Landmark Attention Matters
- Sub-Quadratic Complexity: Standard attention is O(n²); Landmark attention is O(n × k + k²) where k << n — for k = √n, this becomes O(n^1.5), dramatically more efficient for long sequences.
- Global Information Preservation: Unlike local-only attention (which loses distant context), landmark-to-landmark attention maintains a global information pathway — important for tasks requiring full-document understanding.
- Minimal Quality Loss: Well-placed landmarks preserve 95%+ of full attention's information — the compression through landmarks retains the most important global signals.
- Compatible With Flash Attention: The local attention windows and landmark attention patterns can be implemented efficiently with existing optimized kernels.
- Configurable Trade-Off: Adjusting landmark density (k) provides a smooth trade-off between efficiency and information retention — more landmarks = more global information at higher cost.
Landmark Attention Architecture
Landmark Placement Strategies:
- Fixed Stride: Every m-th token is a landmark — simplest, works well for uniform-density text.
- Learned Selection: A scoring network assigns prominence scores; top-k scoring tokens become landmarks — content-aware, better for heterogeneous inputs.
- Boundary-Based: Landmarks placed at sentence boundaries, paragraph breaks, or topic transitions — aligns with natural information structure.
Attention Pattern:
- Regular token t attends to: local window [t−w, t+w] UNION all landmarks.
- Landmark l attends to: its local region UNION all other landmarks.
- This creates a sparse attention pattern with guaranteed global connectivity.
Complexity Comparison
| Method | Attention Complexity | Global Context | Memory |
|---|---|---|---|
| Full Attention | O(n²) | Complete | O(n²) |
| Local Window | O(n × w) | None | O(n × w) |
| Landmark Attention | O(n × k + k²) | Via landmarks | O(n × k) |
| Longformer | O(n × (w + g)) | Via global tokens | O(n × (w + g)) |
Landmark Attention is the information-routing architecture that proves global context can be maintained through strategic compression — using a sparse network of landmark tokens as information hubs that connect distant sequence regions at sub-quadratic cost, achieving the practical efficiency of local attention with the semantic capability of global attention.
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.