Research & Papers

New paper reveals activation patching hides interaction effects in GPT-2 circuits

MIT researchers show interpretability tool actually measures interactions, not just individual components.

Deep Dive

Activation patching is the go-to method for mechanistic interpretability of neural networks, allowing researchers to measure how much each component contributes to a model's output. But a new paper from Vaidyanathan, Arbour, Mueller, Niekum, and Jensen shows this technique has a hidden flaw: the natural indirect effect (NIE) it measures is contaminated by interaction effects (INT) — the component's causal influence depends on the state of other components. By re-deriving the estimand from causal mediation analysis, the authors prove INT is mathematically inevitable and scales with the distance between clean and patched activations.

Demonstrating on the GPT-2 IOI circuit, they find that INT can make some components appear invisible or artificially inflated, and it explains the previously observed instability of faithfulness scores. The paper shows that greedy ranking of components by NIE will miss mechanisms only discoverable through combinatorial search. Crucially, INT is not a nuisance to be eliminated — it serves as a diagnostic: its magnitude and sign reveal when causal conclusions are prompt-dependent. This work challenges standard practices in mechanistic interpretability and suggests researchers need more careful experimental designs.

Key Points
  • Activation patching's NIE estimator inherently includes interaction effects (INT) that scale with activation distance, not just individual component contributions.
  • In GPT-2's IOI circuit, INT caused components to be either invisible or inflated, explaining known instability in faithfulness scores.
  • INT decomposes into pairwise and higher-order group interactions, making greedy ranking by NIE insufficient for discovering all mechanisms.

Why It Matters

Forces AI interpretability researchers to rethink how they measure causal importance, potentially invalidating many past conclusions.

📬 Get the top 10 AI stories daily