arXiv:2609.16487v1 Announce Type: new
Abstract: We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid...
By Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal, Aditya Bansal, Rui Wang, Charles Menguy, Swati Jain
arXiv:2510. 17085v2 Announce Type: replace Abstract: How can we assess the reliability of a dataset without access to ground truth?
By Yiling Chen, Shi Feng, Paul Kattuman, Fang-Yi Yu
arXiv:2608.30842v1 Announce Type: new
Abstract: Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often...
By Deepak Pandita, Christopher M. Homan
The paper introduces a synthetic ground‑truth framework for evaluating explainable AI (XAI) methods, addressing the lack of reliable evaluation procedures. By using controlled interventions to create datasets where the importance of input components is known, the framework generates ground‑truth explanations that align with the model’s actual decision process. The authors apply this approach to binary images, tabular data, and time series, and find that nine popular XAI methods exhibit significant limitations, underscoring the need for intervention‑based benchmarks.
By Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti
The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.
By Timo Freiesleben, Sebastian Zezulka
arXiv:2607. 19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet.
By Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert