arXiv AI

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

arXiv:2607. 01940v1 Announce Type: cross Abstract: Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior.

arXiv Machine Learning
1d ago

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.

By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
arXiv Machine Learning
1d ago

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

The paper investigates the phenomenon of self‑repair in language models, proposing that it arises from a pre‑existing gain in components that act as counterweights when a component is ablated. By modeling interventions as points on a counterfactual axis, the authors derive an affine law for the causal repair response of fine‑grained units, showing that most downstream directions across several models follow this law. They further demonstrate that the slope of this law can be predicted from fixed weights, suggesting that self‑repair is a predictable, counterweight‑driven response rather than a noisy, unexplained effect.

By Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
Hugging Face Trending Papers
Jun 11

Circuit Synchronization Precedes Generalization: Causal Evidence from Fourier Structure in Grokking Transformers

Grokking -- where a transformer on modular arithmetic suddenly transitions from near-chance to near-perfect validation accuracy -- is attributed to a Fourier circuit, but its timing, causal structure, and controllability remain poorly understood. We introduce the Frequency Synchronization Degree (FSD), a normalised, permutation-tested metric for Fourier circuit synchronisation requiring no prior circuit knowledge.