arXiv Machine Learning By Lucia Quirke, Louis Jaburi, David Johnston, William Z. Li, Gon\c{c}alo Paulo, Guillaume Martres, Girish Gupta, Stella Biderman, Nora Belrose

Bergson: An Open Source Library for Data Attribution

Read the original on arXiv Machine Learning →

arXiv:2606. 11660v1 Announce Type: new Abstract: Data attribution is a promising field in interpretability that aims to explain model behavior through the influence of its training data, with applications including debugging undesirable model behavior and training dataset curation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

TrueMuse: A Benchmark for Data Attribution in Text-to-Music Models

TrueMuse is a new benchmark designed to evaluate data attribution in text-to-music models. It consists of a controlled dataset created by fine‑tuning three diffusion‑based models on curated attribution samples, providing known attribution targets. The benchmark covers four settings—melodic structure, timbral characteristics, artist‑level style, and genre‑level patterns—across 133 attributes, 648 models, and 95,456 generated samples, and is used to assess existing black‑box attribution methods along several dimensions.

By Jiawei Yu, Jian Liu
arXiv Machine Learning
Sep 15

Data Attribution at Scale via Influence Matrix Estimation

Data Attribution at Scale via Influence Matrix Estimation proposes a scalable approach to quantify how individual training examples influence a model’s predictions. The authors introduce two algorithms, MAGE and SPELL, that reconstruct an influence matrix from a limited number of measurements without extra computational cost, improving over existing baselines across various training scales and budgets.

By Yuxi Chen, Hamza Golubovic, Han Tong, Arian Maleki, Andrew Ilyas
arXiv AI
Aug 10

MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms

arXiv:2603. 02184v2 Announce Type: replace-cross Abstract: Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mechanisms, has emerged as a promising learning paradigm for conversion rate (CVR) prediction.

By Jinqi Wu, Sishuo Chen, Zhangming Chan, Yong Bai, Lei Zhang, Sheng Chen, Chenghuan Hou, Xiang-Rong Sheng, Han Zhu, Jian Xu, Bo Zheng, Chaoyou Fu