activation patching

**Activation patching** is the **causal intervention method that replaces selected activations in one run with activations from another run to test influence on outputs** - it is one of the most widely used tools in mechanistic interpretability. **What Is Activation patching?** - **Definition**: Patch operation swaps activations at chosen layer, position, and component granularity. - **Purpose**: Measures whether a component carries task-relevant information for target behavior. - **Variants**: Can patch attention head outputs, MLP outputs, residual stream slices, or neuron groups. - **Readout**: Effect size is measured by changes in logits, probabilities, or task success metrics. **Why Activation patching Matters** - **Causal Evidence**: Directly tests necessity and sufficiency of internal signals. - **Circuit Discovery**: Helps isolate components that form behavior-driving pathways. - **Debugging**: Identifies where incorrect behavior first enters computation. - **Safety Analysis**: Useful for tracing risky output generation routes. - **Method Versatility**: Applies across many tasks and model architectures. **How It Is Used in Practice** - **Baseline Design**: Use paired clean and corrupted prompts with clear behavioral contrast. - **Granularity Sweep**: Start broad then narrow to specific heads or features. - **Robustness**: Repeat patch tests across multiple prompt templates to avoid spurious conclusions. Activation patching is **a foundational causal tool for transformer mechanism analysis** - activation patching is most reliable when experiment design cleanly isolates the behavior under study.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account