attention transfer

**Attention Transfer** is a **feature-based knowledge distillation method where the student is trained to mimic the teacher's spatial attention maps** — ensuring the student focuses on the same image regions as the teacher, transferring "what to look at" rather than just "what to predict." **How Does Attention Transfer Work?** - **Attention Map**: $A = sum_c |F_c|^p$ where $F_c$ is the feature map of channel $c$ and $p$ controls the power. - **Loss**: L2 distance between normalized teacher and student attention maps at each layer. - **Layers**: Attention is transferred from multiple intermediate layers simultaneously. - **Paper**: Zagoruyko & Komodakis, "Paying More Attention to Attention" (2017). **Why It Matters** - **Interpretable**: Directly transfers the spatial focus pattern from teacher to student. - **Complementary**: Can be combined with logit-based distillation for stronger knowledge transfer. - **Efficiency**: Small additional computational cost — attention maps are cheap to compute. **Attention Transfer** is **teaching the student where to look** — transferring the teacher's spatial focus patterns to guide the student's feature learning.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account