arXiv:2608. 03629v1 Announce Type: new Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream.
By Abdallah Khemais
The paper introduces OAttention, a token‑level attention mechanism that assigns each token a presence coefficient based on its hidden representation. This coefficient both gates the token’s output and weights its contribution to other tokens, making zero‑vector tokens behave as true zeros and enabling exact null‑receiver, null‑source, and empty‑support properties. The authors extend this idea to local O‑components and an O‑Transformer, and demonstrate small performance changes when retrofitting a pretrained TabPFN model.
By Heyang Gong
The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2606. 21876v2 Announce Type: replace-cross Abstract: The Categorical Jacobian of Zhang et al.
By Rome Thorstenson
arXiv:2605. 23393v2 Announce Type: replace-cross Abstract: Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2508. 08289v3 Announce Type: replace Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave.
By Mu Qiao