arXiv Machine Learning By Rishit Dagli, Abir Harrasse, Luke Zhang, Florent Draye, Amirali Abdullah, Bernhard Sch\"olkopf, Zhijing Jin

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Read the original on arXiv Machine Learning →

arXiv:2606. 05165v1 Announce Type: new Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling

The paper introduces Independent Token Sampling (ITS), a query‑efficient method for detecting memorized training data in diffusion large language models (dLLMs). ITS selects token sets with weak internal dependency by approximating cumulative conditional mutual information using an attention‑derived pairwise dependency proxy and promotes diversity across sampling rounds. Experiments show ITS outperforms existing baselines, improving AUC by 0.18 on the ArXiv dataset while remaining effective under limited query budgets.

By Hongyao Yu, Tianqu Zhuang, Ziyuan Xu, Hao Fang, Jiaxin Hong, Bin Chen, Shu-Tao Xia
arXiv AI
Jul 23

In-Run Data Shapley for Adam Optimizer

arXiv:2602. 00329v4 Announce Type: replace-cross Abstract: Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley value serving as the theoretical gold standard.

By Meng Ding, Zeqing Zhang, Di Wang, Lijie Hu