arXiv Machine Learning By Charlotte H\"ogberg, Ericka Johnson, Kiri L. Wagstaff

Position: Every Ground Truth is a Human Construction, not an Objective Truth

Read the original on arXiv Machine Learning →

arXiv:2607. 09668v1 Announce Type: new Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods

The paper introduces a synthetic ground‑truth framework for evaluating explainable AI (XAI) methods, addressing the lack of reliable evaluation procedures. By using controlled interventions to create datasets where the importance of input components is known, the framework generates ground‑truth explanations that align with the model’s actual decision process. The authors apply this approach to binary images, tabular data, and time series, and find that nine popular XAI methods exhibit significant limitations, underscoring the need for intervention‑based benchmarks.

By Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti
arXiv Statistics ML
2d ago

The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.

By Timo Freiesleben, Sebastian Zezulka