non-local neural networks

**Non-Local Neural Networks** introduce a **non-local operation that captures long-range dependencies in a single layer** — computing the response at each position as a weighted sum of features at all positions, similar to self-attention in transformers but applied to CNNs. **How Do Non-Local Blocks Work?** - **Formula**: $y_i = frac{1}{C(x)} sum_j f(x_i, x_j) cdot g(x_j)$ - **$f$**: Pairwise affinity function (embedded Gaussian, dot product, or concatenation). - **$g$**: Value transformation (linear embedding). - **Residual**: $z_i = W_z y_i + x_i$ (residual connection). - **Paper**: Wang et al. (2018). **Why It Matters** - **Long-Range**: Captures dependencies between distant positions in a single layer (vs. CNN's local receptive field). - **Video**: Particularly effective for video understanding where temporal long-range dependencies are critical. - **Pre-ViT**: Brought self-attention to computer vision before Vision Transformers existed. **Non-Local Networks** are **self-attention for CNNs** — the bridge concept that brought transformer-style global interaction to convolutional architectures.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account