The paper introduces two algorithms, MAGE and SPELL, that enable efficient data attribution in neural networks by estimating a large influence matrix from a limited number of measurements. These methods leverage existing metagradient techniques without additional computational overhead, addressing the challenge of predicting the impact of removing training data in non‑convex models. Experiments show that MAGE and SPELL outperform current baselines across various training scales and measurement budgets.
By Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas
arXiv:2402. 08922v3 Announce Type: replace Abstract: Large-scale black-box models have become ubiquitous across numerous applications.
By Myeongseob Ko, Feiyang Kang, Weiyan Shi, Ming Jin, Zhou Yu, Ruoxi Jia
arXiv:2606. 05165v1 Announce Type: new Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data.
By Rishit Dagli, Abir Harrasse, Luke Zhang, Florent Draye, Amirali Abdullah, Bernhard Sch\"olkopf, Zhijing Jin
arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.
By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu
arXiv:2501.14271v4 Announce Type: replace
Abstract: Meta-learning enables models to rapidly adapt to new tasks by leveraging prior experience, but its adaptation mechanisms remain opaque, especially...
By Yoshihiro Mitsuka, Shadan Golestan, Zahin Sufiyan, Shotaro Miwa, Osmar R. Zaiane
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer