Demystifying Variance in Circuit Discovery of LLMs
arXiv:2606. 16920v1 Announce Type: cross Abstract: Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task.
The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.
arXiv:2606. 16920v1 Announce Type: cross Abstract: Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a given task.
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, demonstrating that accurate average estimates can still lead to poor decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B show that observers trained on action loss tend to select lower‑loss actions, while traditional metrics like AUROC can rank monitors differently from deployment loss, highlighting the need for task‑specific evaluation.
arXiv:2601. 09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models.
arXiv:2607. 18921v1 Announce Type: cross Abstract: Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior.
arXiv:2609.36813v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong general capabilities that mechanistic interpretability has attributed to sparse computational circuits. How...
The paper investigates whether structural differences in circuits discovered by circuit discovery methods reflect distinct mechanisms. By varying input-token frequency while keeping the task constant, the authors find that although circuits appear specialized by frequency structurally, functional and representational analyses reveal no reliable differences, a phenomenon they call phantom specialization. Across multiple models and tasks, structurally distinct circuits implement the same computation, with core shared subgraphs recovering most of the performance and interchangeable internal representations confirmed by causal interventions.
The study investigates the consistency and specificity of language model circuits across six tasks and five models, focusing on component-level (attention heads and MLP blocks) and neuron-level circuits. Component-level circuits are highly consistent and causally important but lack task specificity, as ablating a circuit for one task similarly harms performance on other tasks. Neuron-level circuits show higher task specificity but lower consistency, with overlap mainly between closely related tasks. The analysis of Llama‑3.2‑3B reveals that shared components are predominantly MLP blocks, while attention heads act as generic attention‑sink heads.
The paper introduces Circuit Condensation, a post‑training method that prunes low‑attribution edges from large causal graphs and trains a low‑rank adapter to preserve behavior. Across four behaviors and eight models, the condensed circuits are on average 8.1× smaller than the strongest frozen baseline, with reductions up to 316×. Experiments show that weight updates drive the size reduction, and detailed ablations reveal dependencies among remaining edges and a more focused set of heads for indirect object identification.
arXiv:2608. 13754v1 Announce Type: new Abstract: The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions.
arXiv:2606. 24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize.
The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.