arXiv Machine Learning

Attribution via Distributional Paths for Information Revelation

arXiv:2606. 03885v1 Announce Type: new Abstract: Feature attribution methods explain predictions by assigning importance scores to input features.

arXiv Machine Learning
Jun 10

XtrAIn: Training-Guided Occlusion for Feature Attribution

arXiv:2606. 10877v1 Announce Type: new Abstract: Occlusion-based attribution methods provide an intuitive way to estimate feature importance by perturbing input features and measuring the resulting change in model output.

By Thodoris Lymperopoulos, Ioannis Kakogeorgiou, Denia Kanellopoulou
arXiv Machine Learning
Aug 28

The Attribution Contract for Generative Language Models

The paper argues that feature attribution scores for generative language models lack a fixed meaning because each generated token is both output and input, leading to multiple distinct explanatory questions. It introduces the Attribution Contract framework, which explicitly defines the model score, fixed variables, target output, generation process, and eligible features, showing how these choices affect attribution outcomes. Experiments demonstrate that different contracts (e.g., local next-token vs. prompt-level) and model architectures (mixture-of-experts vs. masked-diffusion) yield markedly different attribution distributions, highlighting the need for careful contract specification.

By Giang Nguyen
arXiv Machine Learning
Aug 31

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

The paper introduces Concept-Targeted Attribution (CTA), a method that trains attribution graphs to explain the emergence of internal concept representations in language models, rather than just the final token prediction. CTA produces probe-specific circuits that reveal which internal computations drive a linear probe’s accuracy, and cross-layer transcoders demonstrate that these graphs contain predictive structure across multiple concept categories. Causal ablations show that probe-targeted and logit-targeted graphs capture distinct mechanisms, with probe-relevant features affecting internal concept scores and logit-relevant features altering generated tokens.

By Vedant Palit, Florent Draye, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin
arXiv AI
Sep 25

Every Component Is a Lookup: One Linear Graph for Interaction, Composition and Attribution

The paper proposes that two architectural assumptions—(1) attention and MLPs share a key‑value form <phi(S)>U, and (2) components read from an additive residual stream—are sufficient to answer three interpretability questions: component interaction, information routing, and token attribution. By treating these selections as a computational graph, the authors develop Unpack, a backward attribution method that validates interaction scores, recovered routes, and token attribution against established tests across models ranging from 160M to 6.9B parameters. The study also shows that contribution and causal effect can differ, with a recognizable signature in how components change when a task is removed.

By Po-Kai Chen, Aske Plaat, Niki van Stein
arXiv Machine Learning
Sep 18

Evaluating Explanation Methods by the Predictors They Induce

The paper proposes a new evaluation test for explanation methods: if an explanation accurately captures how a model uses its features, one should be able to reconstruct the model’s predictions from it. The authors convert explanations into predictors by summing feature effects and assess how well these predictors reproduce the model on unseen data, without any fitting. They apply this test to partial dependence plots, accumulated local effects, SHAP, and LIME across multiple datasets and model families, showing that the best method depends on feature dependence and that some existing quality metrics can favor flawed explanations.

By Jacob Selb{\ae}k, Hugo L. Hammer
arXiv Computer Vision
Sep 23

MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models

MorphoSHAP is a model‑agnostic post‑hoc explanation method that uses morphological shapes—derived from the Tree of Shapes—as the units of attribution in a Shapley game. Each shape is characterized by its scale, geometry, and signed contribution, enabling explanations that reveal where evidence lies, what structural type carries it, and how strongly it influences the prediction. The approach offers spatial, textual, and global class‑level explanations, surpasses existing attribution methods on multiple benchmarks, and is preferred by users in a study.

By Anirudh Prabhakaran, Alexandre Rocchi, Gianni Franchi
arXiv AI
Sep 16

ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers

The paper introduces ResLRP, an extension of Layer-wise Relevance Propagation that explicitly handles residual connections in Vision Transformers to prevent attribution explosion. It demonstrates that residual cancellation causes instability in ViT explanations, and that ResLRP improves faithfulness and localization across a wide range of ViT architectures, including Vision Language Models. The method also provides a diagnostic measure for predicting attribution degradation and successfully localizes Sparse Autoencoder features.

By Jim Berend, Reduan Achtibat, Daniel Sch\"affer, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer
arXiv Machine Learning
Jul 1

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

arXiv:2606. 32008v1 Announce Type: new Abstract: Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens.

By Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo, Chia-Tse Shao, Yingxiao Ye, Aobo Yang, Vivek Miglani, Nehal Bandi