sandwich transformer
**Sandwich Transformer** is a **transformer variant that reorders self-attention and feedforward sublayers** — placing attention sublayers in the middle of the network and feedforward sublayers at the top and bottom, creating a "sandwich" structure that improves perplexity.
**How Does Sandwich Transformer Work?**
- **Standard Transformer**: Alternating [Attention, FFN, Attention, FFN, ...].
- **Sandwich**: [FFN, FFN, ..., Attention, Attention, ..., FFN, FFN, ...].
- **Reordering**: Attention layers are concentrated in the middle, FFN layers at the boundaries.
- **Paper**: Press et al. (2020).
**Why It Matters**
- **Free Improvement**: Simply reordering sublayers (no new parameters) improves language modeling perplexity.
- **Insight**: Suggests that the standard alternating pattern may not be optimal.
- **Architecture Search**: Motivates searching over sublayer orderings, not just sublayer types.
**Sandwich Transformer** is **transformer with rearranged layers** — the surprising finding that putting attention in the middle and FFN at the edges improves performance for free.