arXiv:2606. 05165v1 Announce Type: new Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data.
By Rishit Dagli, Abir Harrasse, Luke Zhang, Florent Draye, Amirali Abdullah, Bernhard Sch\"olkopf, Zhijing Jin
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
By Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
The paper introduces dattri-LLM, a library designed to make training data attribution (TDA) practical for large language models. It achieves efficiency by using compact gradient representations and a cost‑based routing system, while maintaining compatibility by capturing per‑example gradients from existing training loops without modifications, even in distributed settings. The library also offers extensibility through reusable gradient operations and callbacks, supporting various attribution methods and applications such as online data selection, and demonstrates significant performance gains and scalability up to 110B‑parameter models.
By Shixuan Liu, Tongli Zhou, Junwei Deng, Pingbang Hu, Jiaqi W. Ma
arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).
By Steffen Dereich, Arnulf Jentzen, Adrian Riekert
The paper introduces two algorithms, MAGE and SPELL, that enable efficient data attribution in neural networks by estimating a large influence matrix from a limited number of measurements. These methods leverage existing metagradient techniques without additional computational overhead, addressing the challenge of predicting the impact of removing training data in non‑convex models. Experiments show that MAGE and SPELL outperform current baselines across various training scales and measurement budgets.
By Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas
arXiv:2607. 21094v1 Announce Type: new Abstract: We study feature-level and node-level explanations for graph neural networks (GNNs) through the lens of Aumann-Shapley attribution.
By Bizu Feng, Zhimu Yang, Shuming Wang, Shaode Yu, Yuan Cheng, Xiaojun Qian, Zixin Hu