arXiv Machine Learning

Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs

arXiv Machine Learning
Aug 31

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

The paper introduces Concept-Targeted Attribution (CTA), a method that trains attribution graphs to explain the emergence of internal concept representations in language models, rather than just the final token prediction. CTA produces probe-specific circuits that reveal which internal computations drive a linear probe’s accuracy, and cross-layer transcoders demonstrate that these graphs contain predictive structure across multiple concept categories. Causal ablations show that probe-targeted and logit-targeted graphs capture distinct mechanisms, with probe-relevant features affecting internal concept scores and logit-relevant features altering generated tokens.

By Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin
arXiv Machine Learning
Aug 5

LLMs Can Annotate Attribution Graphs

arXiv:2608. 02632v1 Announce Type: new Abstract: Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes.

By Ameen Patel, Max Zhang, Nathan Hu
arXiv Machine Learning
Aug 11

Can Graph Learning Learn Circuits?

arXiv:2608. 08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient to reproduce a particular behavior.

By Chester Tan, Moritz Lampert, Courtney Maynard, Ankit Ramakrishnan, Tina Eliassi-Rad, Ingo Scholtes
arXiv Machine Learning
Sep 17

ReDIL-GNN: Resynthesis Domain Incremental Learning for Circuit Graph Neural Networks

ReDIL-GNN is a framework for resynthesis domain‑incremental learning in circuit graph neural networks. It adapts a fixed prediction or representation head as new synthesis styles appear and evaluates retention across all previously seen domains. The method introduces the Resynthesis Adaptability Index (RAI), a pre‑adaptation score that combines adaptation need, source‑equivalence recoverability, structural coverage, and update compatibility to decide whether to adapt, reuse, or defer updates.

By Rupesh Raj Karn, Johann Knechtel, Ozgur Sinanoglu
arXiv AI
Aug 28

Diff Mining: Logit Differences Reveal Finetuning Objectives

Diff Mining is a framework that identifies what a finetuned language model has learned by comparing its logits to those of its base model. It extracts per-context logit differences on a reference corpus and aggregates them into an interpretable token set using either a Top‑K frequency method or Non‑negative Matrix Factorization. The approach outperforms existing model‑diffing methods in domain detection and bias identification, and it requires only access to output logits, making it scalable to large models.

By Greg Kocher, Robert West, Cl\'ement Dumas, Julian Minder