arXiv AI

Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes

arXiv:2606. 09607v1 Announce Type: cross Abstract: Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clustering co-activation statistics.

arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv AI
Jun 2

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

arXiv:2606. 02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (dense transformer, mixture-of-experts) and two pretraining corpora (The Pile, DCLM): Pythia 1B, OLMo 1B-0724-hf, and OLMoE 1B-7B-0924.

By Yongzhong Xu
arXiv Computation and Language
Aug 28

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

The paper investigates whether structural differences in circuits discovered by circuit discovery methods reflect distinct mechanisms. By varying input-token frequency while keeping the task constant, the authors find that although circuits appear specialized by frequency structurally, functional and representational analyses reveal no reliable differences, a phenomenon they call phantom specialization. Across multiple models and tasks, structurally distinct circuits implement the same computation, with core shared subgraphs recovering most of the performance and interchangeable internal representations confirmed by causal interventions.

By Alireza Bayat Makou, Jingcheng Niu, Subhabrata Dutta, Iryna Gurevych
arXiv Machine Learning
Jun 2

Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation

arXiv:2505. 17630v4 Announce Type: replace-cross Abstract: Circuit localization methods aim to identify the subset of model components responsible for specific behaviors in large language models, enabling detailed mechanistic analysis.

By Joakim Edin, Casper L. Christensen, R\'obert Csord\'as, Tuukka Ruotsalo, Zhengxuan Wu, Maria Maistro, Jing Huang, Lars Maal{\o}e
arXiv Machine Learning
Aug 28

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit

The paper introduces Circuit Condensation, a post‑training method that prunes low‑attribution edges from large causal graphs and trains a low‑rank adapter to preserve behavior. Across four behaviors and eight models, the condensed circuits are on average 8.1× smaller than the strongest frozen baseline, with reductions up to 316×. Experiments show that weight updates drive the size reduction, and detailed ablations reveal dependencies among remaining edges and a more focused set of heads for indirect object identification.

By Sai Adith Senthil Kumar
arXiv AI
Sep 10

Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

The paper investigates whether recent attention‑mechanism improvements—specifically gated attention, Kimi K3, Kimi Delta Attention, and Attention Residuals—effectively eliminate the attention‑sink problem when scaling language models to a one‑million‑token context window. Using a new diagnostic suite called SinkProbe, the authors evaluate sink mass, massive activation, position‑resolved recall, and the recency gap across four small models that vary only in token mixing and depth. Their findings show that the training objective, rather than the architecture, drives the emergence of attention sinks; gating did not replicate its previously reported benefits at the larger scale, and sink mass, activations, and positional bias behaved independently.

By Sara Rizwan, Samaanah Abdus Salam
arXiv Computation and Language
Sep 4

How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits

The study investigates the consistency and specificity of language model circuits across six tasks and five models, focusing on component-level (attention heads and MLP blocks) and neuron-level circuits. Component-level circuits are highly consistent and causally important but lack task specificity, as ablating a circuit for one task similarly harms performance on other tasks. Neuron-level circuits show higher task specificity but lower consistency, with overlap mainly between closely related tasks. The analysis of Llama‑3.2‑3B reveals that shared components are predominantly MLP blocks, while attention heads act as generic attention‑sink heads.

By Michael Li, Nishant Subramani