induction heads
**Induction heads** is the **attention heads that implement next-token continuation by matching repeated token patterns in context** - they are a canonical example of interpretable in-context learning circuitry.
**What Is Induction heads?**
- **Definition**: Head pattern often attends from a repeated token to the token that followed its prior occurrence.
- **Functional Role**: Supports copying and continuation behavior after seeing a short pattern once.
- **Layer Pattern**: Usually appears in mid-to-late layers where richer context features exist.
- **Circuit Context**: Often works with earlier heads that mark previous-token relationships.
**Why Induction heads Matters**
- **Interpretability Landmark**: Provides a concrete, testable mechanism for in-context behavior.
- **Generalization Insight**: Shows how transformers can implement algorithm-like pattern reuse.
- **Safety Relevance**: Helps explain unintended copying and memorization pathways.
- **Model Comparison**: Useful benchmark for checking mechanism emergence across scales.
- **Tool Validation**: Frequently used to evaluate causal interpretability methods.
**How It Is Used in Practice**
- **Prompt Probes**: Use synthetic repeated-pattern prompts to isolate induction behavior.
- **Head Patching**: Patch candidate head activations to verify continuation dependence.
- **Ablation Checks**: Disable candidate heads and measure drop in pattern-continuation accuracy.
Induction heads is **a well-studied mechanistic motif in transformer attention** - induction heads remain a key reference mechanism for connecting attention structure to concrete behavior.