The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
By Po-Kai Chen, Aske Plaat, Niki van Stein
The paper investigates how many transformer components influence a token prediction by measuring the absolute contribution of each unit and channel to the logit. It finds that thousands of components contribute to a single prediction, yet a small subset—often just dozens—carries the majority of the predictive mass. Across models ranging from 124 M to 7 B parameters, the proportion of the model involved in a prediction remains around one to three percent, independent of size, and the study demonstrates that specific components can be directly read and written to modify model behavior without additional training.
By Mark Oskin
The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.
By Richard Zhe Wang
Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.
By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts
The paper argues that feature attribution scores for generative language models lack a fixed meaning because each generated token is both output and input, leading to multiple distinct explanatory questions. It introduces the Attribution Contract framework, which explicitly defines the model score, fixed variables, target output, generation process, and eligible features, showing how these choices affect attribution outcomes. Experiments demonstrate that different contracts (e.g., local next-token vs. prompt-level) and model architectures (mixture-of-experts vs. masked-diffusion) yield markedly different attribution distributions, highlighting the need for careful contract specification.
By Giang Nguyen
arXiv:2608. 03629v1 Announce Type: new Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream.
By Abdallah Khemais