The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.
By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
The paper investigates the phenomenon of self‑repair in language models, proposing that it arises from a pre‑existing gain in components that act as counterweights when a component is ablated. By modeling interventions as points on a counterfactual axis, the authors derive an affine law for the causal repair response of fine‑grained units, showing that most downstream directions across several models follow this law. They further demonstrate that the slope of this law can be predicted from fixed weights, suggesting that self‑repair is a predictable, counterweight‑driven response rather than a noisy, unexplained effect.
By Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
arXiv:2607. 18921v1 Announce Type: cross Abstract: Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior.
By Yang Sheng, Jie Fu
arXiv:2608. 14689v1 Announce Type: new Abstract: Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently.
By Pratikkumar Babariya
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
By Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen
arXiv:2609.30465v1 Announce Type: cross
Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed...
By Mingyang Song, Mao Zheng
arXiv:2607. 27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing.
By Lei Dong
arXiv:2609.36813v1 Announce Type: new
Abstract: Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. How...
By Chuanpu Liu, Miao Yu, Yikai Cai, Yuanhe Zhang, Zhenhong Zhou, Li Sun, Zuming Jiang, Yufei Guo
arXiv:2605. 23393v2 Announce Type: replace-cross Abstract: Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computational route that produced a prediction.
By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv:2608. 15286v1 Announce Type: cross Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym.
By Shiven Khurdi
Grokking -- where a transformer on modular arithmetic suddenly transitions from near-chance to near-perfect validation accuracy -- is attributed to a Fourier circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior circuit knowledge.
arXiv:2608. 03620v1 Announce Type: cross Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass.
By Abdallah Khemais