Home Knowledge Base Twins Transformer

Twins Transformer is a hierarchical vision Transformer that introduces spatially separable self-attention (SSSA), combining local attention within sub-windows with global attention through sub-sampled key-value tokens, achieving efficient multi-scale feature extraction with both fine-grained local and coarse global spatial interactions. Twins comes in two variants: Twins-PCPVT (using conditional position encoding from PVT) and Twins-SVT (using spatially separable attention).

Why Twins Transformer Matters in AI/ML: Twins Transformer provides efficient global-local attention that captures both fine-grained local patterns and global context without the quadratic cost of full attention, achieving strong performance on classification, detection, and segmentation with a simple, elegant design.

Locally-Grouped Self-Attention (LSA) — The feature map is divided into non-overlapping sub-windows (similar to Swin), and self-attention is computed independently within each sub-window at O(N·w²) cost; this captures detailed local interactions efficiently • Global Sub-Sampled Attention (GSA) — A single representative token is extracted from each sub-window (via average pooling or learned aggregation), and global attention is computed among these representative tokens; the result is broadcast back to all tokens, providing global context at O(N·(N/w²)) cost • Alternating LSA and GSA — Twins-SVT alternates between LSA layers (local attention within windows) and GSA layers (global attention via sub-sampling), ensuring every token eventually interacts with every other token through the combination of local and global mechanisms • Conditional Position Encoding (CPE) — Twins-PCPVT uses depth-wise convolutions as position encoding (applied after each attention layer), eliminating fixed or learned position embeddings and enabling variable input resolutions without interpolation • Hierarchical design — Like PVT and Swin, Twins uses a 4-stage pyramidal architecture with progressive spatial downsampling, producing multi-scale features compatible with FPN-based detection and segmentation heads

Attention TypeScopeComplexityRole
LSA (Local)Within sub-windowsO(N·w²)Fine-grained local patterns
GSA (Global)Sub-sampled globalO(N·N/w²)Global context aggregation
CombinedFull coverageO(N·(w² + N/w²))Local detail + global context
Swin (comparison)Shifted windowsO(N·w²)Local with shift-based global
PVT SRA (comparison)Reduced keys/valuesO(N·N/R²)Full attention, reduced cost

Twins Transformer provides an elegant solution to the local-global attention tradeoff through spatially separable self-attention, alternating efficient local window attention with sub-sampled global attention to achieve comprehensive spatial coverage at sub-quadratic cost, establishing a powerful design principle for efficient hierarchical vision Transformers.

twins transformercomputer vision

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.