arXiv Machine Learning By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Read the original on arXiv Machine Learning →

Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 31

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

The paper introduces Concept-Targeted Attribution (CTA), a method that trains attribution graphs to explain the emergence of internal concept representations in language models, rather than just the final token prediction. CTA produces probe-specific circuits that reveal which internal computations drive a linear probe’s accuracy, and cross-layer transcoders demonstrate that these graphs contain predictive structure across multiple concept categories. Causal ablations show that probe-targeted and logit-targeted graphs capture distinct mechanisms, with probe-relevant features affecting internal concept scores and logit-relevant features altering generated tokens.

By Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin