Sandwich Transformer is a transformer variant that reorders self-attention and feedforward sublayers — placing attention sublayers in the middle of the network and feedforward sublayers at the top and bottom, creating a "sandwich" structure that improves perplexity.
How Does Sandwich Transformer Work?
- Standard Transformer: Alternating [Attention, FFN, Attention, FFN, ...].
- Sandwich: [FFN, FFN, ..., Attention, Attention, ..., FFN, FFN, ...].
- Reordering: Attention layers are concentrated in the middle, FFN layers at the boundaries.
- Paper: Press et al. (2020).
Why It Matters
- Free Improvement: Simply reordering sublayers (no new parameters) improves language modeling perplexity.
- Insight: Suggests that the standard alternating pattern may not be optimal.
- Architecture Search: Motivates searching over sublayer orderings, not just sublayer types.
Sandwich Transformer is transformer with rearranged layers — the surprising finding that putting attention in the middle and FFN at the edges improves performance for free.
sandwich transformerefficient transformer
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.