arXiv:2609.16487v1 Announce Type: new
Abstract: We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid...
By Aniruddha Tamhane, Raghavendra Addanki, Ayushi Aggarwal, Aditya Bansal, Rui Wang, Charles Menguy, Swati Jain
arXiv:2510. 17085v2 Announce Type: replace Abstract: How can we assess the reliability of a dataset without access to ground truth?
By Yiling Chen, Shi Feng, Paul Kattuman, Fang-Yi Yu
arXiv:2608.30842v1 Announce Type: new
Abstract: Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often...
By Deepak Pandita, Christopher M. Homan
The paper introduces a synthetic ground‑truth framework for evaluating explainable AI (XAI) methods, addressing the lack of reliable evaluation procedures. By using controlled interventions to create datasets where the importance of input components is known, the framework generates ground‑truth explanations that align with the model’s actual decision process. The authors apply this approach to binary images, tabular data, and time series, and find that nine popular XAI methods exhibit significant limitations, underscoring the need for intervention‑based benchmarks.
By Miquel Mir\'o-Nicolau, Francesco Spinnato, Riccardo Guidotti
The article discusses how predictive benchmarking—evaluating machine learning models by their performance and ranking—serves as a core method in machine learning research. It argues that benchmark scores only reflect performance on specific datasets and learning problems, and that drawing broader scientific conclusions requires explicit assumptions. By adapting concepts from psychological validity theory, the authors propose validity conditions to make these assumptions clear, and demonstrate their application in two case studies (ImageNet and the Fragile Families Challenge) to illustrate how benchmark results can inform inferences about research progress and limits of predictability.
By Timo Freiesleben, Sebastian Zezulka
arXiv:2607. 19355v1 Announce Type: new Abstract: LLMs are increasingly used with external knowledge sources like the internet.
By Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, Eric Gilbert
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
By Arthur Hoarau
arXiv:2607. 09489v1 Announce Type: new Abstract: An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation.
By Jade Alglave, Patrick Cousot
arXiv:2408. 02379v2 Announce Type: replace-cross Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act.
By Benjamin Fresz, Vincent Philipp G\"obels, Safa Omri, Danilo Brajovic, Andreas Aichele, Janika Kutz, Jens Neuh\"uttler, Marco F. Huber
The article "Beyond RAGs: Building Actually Truthful AI Harnesses" discusses the limitations of Retrieval-Augmented Generation (RAG) systems, emphasizing that retrieval alone does not guarantee evidence for AI claims. It explores methods for constructing AI systems that can substantiate their statements, moving beyond simple retrieval to more robust proof mechanisms. The piece highlights the importance of developing AI that can verify its own outputs rather than merely retrieve information.
By Ari Joury, PhD
arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
By Ranjan Mishra, Jakob Schoeffer
This PhD thesis examines how machine learning (ML) influences society, noting that ML increasingly shapes consequential decisions and recommendations. It highlights the risk of discriminatory effects when fairness is not explicitly considered in data‑driven systems. The work proposes methods for measuring fairness, decomposing ML systems to anticipate bias, and implementing interventions that reduce discrimination while preserving utility, and it outlines future research directions as ML, including generative AI, becomes more integrated into society.
By Joachim Baumann