Non-Local Neural Networks introduce a non-local operation that captures long-range dependencies in a single layer — computing the response at each position as a weighted sum of features at all positions, similar to self-attention in transformers but applied to CNNs.
How Do Non-Local Blocks Work?
- Formula: $y_i = frac{1}{C(x)} sum_j f(x_i, x_j) cdot g(x_j)$
- $f$: Pairwise affinity function (embedded Gaussian, dot product, or concatenation).
- $g$: Value transformation (linear embedding).
- Residual: $z_i = W_z y_i + x_i$ (residual connection).
- Paper: Wang et al. (2018).
Why It Matters
- Long-Range: Captures dependencies between distant positions in a single layer (vs. CNN's local receptive field).
- Video: Particularly effective for video understanding where temporal long-range dependencies are critical.
- Pre-ViT: Brought self-attention to computer vision before Vision Transformers existed.
Non-Local Networks are self-attention for CNNs — the bridge concept that brought transformer-style global interaction to convolutional architectures.
non-local neural networkscomputer vision
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.