arXiv Machine Learning

Position: Every Ground Truth is a Human Construction, not an Objective Truth

arXiv:2607. 09668v1 Announce Type: new Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.

arXiv AI
6d ago

A Synthetic Ground-Truth Framework for the Evaluation of Explainable AI Methods

The paper introduces a synthetic ground‑truth framework for evaluating explainable AI (XAI) methods, addressing the lack of reliable evaluation procedures. By using controlled interventions to create datasets where the importance of input components is known, the framework generates ground‑truth explanations that align with the model’s actual decision process. The authors apply this approach to binary images, tabular data, and time series, and find that nine popular XAI methods exhibit significant limitations, underscoring the need for intervention‑based benchmarks.

By Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti
arXiv Statistics ML
2d ago

The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.

By Timo Freiesleben, Sebastian Zezulka
arXiv AI
Jul 10

The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis

arXiv:2408. 02379v2 Announce Type: replace-cross Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act.

By Benjamin Fresz, Vincent Philipp G\"obels, Safa Omri, Danilo Brajovic, Andreas Aichele, Janika Kutz, Jens Neuh\"uttler, Marco F. Huber
Towards Data Science
Sep 24

Beyond RAGs: Building Actually Truthful AI Harnesses

The article "Beyond RAGs: Building Actually Truthful AI Harnesses" discusses the limitations of Retrieval-Augmented Generation (RAG) systems, emphasizing that retrieval alone does not guarantee evidence for AI claims. It explores methods for constructing AI systems that can substantiate their statements, moving beyond simple retrieval to more robust proof mechanisms. The piece highlights the importance of developing AI that can verify its own outputs rather than merely retrieve information.

By Ari Joury, PhD
arXiv Machine Learning
Sep 11

On the Societal Impact of Machine Learning

This PhD thesis examines how machine learning (ML) influences society, noting that ML increasingly shapes consequential decisions and recommendations. It highlights the risk of discriminatory effects when fairness is not explicitly considered in data‑driven systems. The work proposes methods for measuring fairness, decomposing ML systems to anticipate bias, and implementing interventions that reduce discrimination while preserving utility, and it outlines future research directions as ML, including generative AI, becomes more integrated into society.

By Joachim Baumann