arXiv AI By Zhiren Gong, Zihao Zeng, Chau Yuen, Wei Yang Bryan Lim

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

Read the original on arXiv AI →

arXiv:2607. 01940v1 Announce Type: cross Abstract: Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.

By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
arXiv Machine Learning
1d ago

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

The paper investigates the phenomenon of self‑repair in language models, proposing that it arises from a pre‑existing gain in components that act as counterweights when a component is ablated. By modeling interventions as points on a counterfactual axis, the authors derive an affine law for the causal repair response of fine‑grained units, showing that most downstream directions across several models follow this law. They further demonstrate that the slope of this law can be predicted from fixed weights, suggesting that self‑repair is a predictable, counterweight‑driven response rather than a noisy, unexplained effect.

By Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu