arXiv Machine Learning

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

arXiv:2606. 05165v1 Announce Type: new Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data.

arXiv Machine Learning
Sep 22

Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling

The paper introduces Independent Token Sampling (ITS), a query‑efficient method for detecting memorized training data in diffusion large language models (dLLMs). ITS selects token sets with weak internal dependency by approximating cumulative conditional mutual information using an attention‑derived pairwise dependency proxy and promotes diversity across sampling rounds. Experiments show ITS outperforms existing baselines, improving AUC by 0.18 on the ArXiv dataset while remaining effective under limited query budgets.

By Hongyao Yu, Tianqu Zhuang, Ziyuan Xu, Hao Fang, Jiaxin Hong, Bin Chen, Shu-Tao Xia
arXiv AI
Jul 23

In-Run Data Shapley for Adam Optimizer

arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.

By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu
arXiv Machine Learning
Jun 18

BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training

arXiv:2606. 18650v1 Announce Type: new Abstract: As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories.

By Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia Jiang
arXiv AI
3d ago

dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale

The paper introduces dattri-LLM, a library designed to make training data attribution (TDA) practical for large language models. It achieves efficiency by using compact gradient representations and a cost‑based routing system, while maintaining compatibility by capturing per‑example gradients from existing training loops without modifications, even in distributed settings. The library also offers extensibility through reusable gradient operations and callbacks, supporting various attribution methods and applications such as online data selection, and demonstrates significant performance gains and scalability up to 110B‑parameter models.

By Shixuan Liu, Tongli Zhou, Junwei Deng, Pingbang Hu, Jiaqi W. Ma