arXiv Machine Learning

Data Attribution at Scale via Influence Matrix Estimation

Data Attribution at Scale via Influence Matrix Estimation proposes a scalable approach to quantify how individual training examples influence a model’s predictions. The authors introduce two algorithms, MAGE and SPELL, that reconstruct an influence matrix from a limited number of measurements without extra computational cost, improving over existing baselines across various training scales and budgets.

arXiv Machine Learning
Sep 21

Data Attribution via Sketched Metadifferentiation

The paper introduces two algorithms, MAGE and SPELL, that enable efficient data attribution in neural networks by estimating a large influence matrix from a limited number of measurements. These methods leverage existing metagradient techniques without additional computational overhead, addressing the challenge of predicting the impact of removing training data in non‑convex models. Experiments show that MAGE and SPELL outperform current baselines across various training scales and measurement budgets.

By Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas
arXiv AI
Jul 23

In-Run Data Shapley for Adam Optimizer

arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.

By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu
arXiv Machine Learning
Jun 11

Bergson: An Open Source Library for Data Attribution

arXiv:2606. 11660v1 Announce Type: new Abstract: Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation.

By Lucia Quirke, Louis Jaburi, David Johnston, William Z. Li, Gon\c{c}alo Paulo, Guillaume Martres, Girish Gupta, Stella Biderman, Nora Belrose
arXiv Machine Learning
Sep 23

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.

By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts
arXiv AI
3d ago

Distilling Diffusion Score Discrepancy for Efficient Training Data Attribution

The paper introduces a new method for training data attribution in diffusion models called TID, which uses a local score discrepancy measure and can be estimated without retraining. It further distills this approach into TIDE, a forward‑only student that reproduces the teacher’s rankings using internal activations, achieving comparable accuracy at dramatically lower query cost. Experiments on CIFAR‑10, ArtBench‑10, and MS‑COCO show that TID outperforms existing methods and TIDE attributes samples in milliseconds, faster than generation itself.

By Shixuan Liu, Joan Serr\`a, Kin Wai Cheuk, Jinju Kim, Woosung Choi, Yukara Ikemiya, Wei-Hsiang Liao, Jiaqi W. Ma, Yuki Mitsufuji
arXiv AI
3d ago

dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale

The paper introduces dattri-LLM, a library designed to make training data attribution (TDA) practical for large language models. It achieves efficiency by using compact gradient representations and a cost‑based routing system, while maintaining compatibility by capturing per‑example gradients from existing training loops without modifications, even in distributed settings. The library also offers extensibility through reusable gradient operations and callbacks, supporting various attribution methods and applications such as online data selection, and demonstrates significant performance gains and scalability up to 110B‑parameter models.

By Shixuan Liu, Tongli Zhou, Junwei Deng, Pingbang Hu, Jiaqi W. Ma
arXiv Machine Learning
Jul 31

What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis

arXiv:2510. 03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the machine learning community.

By Shahriar Kabir Nahin, Wenxiao Xiao, Joshua Liu, Anshuman Chhabra, Hongfu Liu