CausalProfiler is a synthetic benchmark generator designed to evaluate causal machine learning (Causal ML) methods more rigorously and transparently. It randomly samples causal models, data, queries, and ground truths based on explicit design choices across observation, intervention, and counterfactual reasoning levels, providing coverage guarantees and transparent assumptions. The authors demonstrate its utility by testing several state‑of‑the‑art methods under diverse conditions, both within and outside the identification regime, highlighting the insights CausalProfiler can reveal.
By Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, \"Ozg\"ur \c{S}im\c{s}ek
arXiv:2605. 05882v2 Announce Type: replace-cross Abstract: Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protected attributes such as race, gender, or age.
By Filip Edstr\"om, Guilherme W. F. Barros, Tetiana Gorbach, Xavier de Luna
arXiv:2607. 07762v1 Announce Type: new Abstract: Modern machine learning (ML) increasingly relies on complex models whose behavior is difficult to characterize beyond empirical performance metrics.
By Thibaut Vidal, Julien Ferry
Modern machine learning (ML) increasingly relies on complex models whose behavior is difficult to characterize beyond empirical performance metrics. Across a wide range of tasks, including prediction, generation, and decision-making, models with similar empirical performance can exhibit markedly different properties in terms of their transparency, interpretability, robustness, fairness, privacy, and certifiability.
arXiv:2608. 02238v1 Announce Type: cross Abstract: Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health.
By Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh
arXiv:2511.21799v2 Announce Type: replace
Abstract: Real-world machine learning (ML) pipelines rarely produce a single model; instead, they produce a Rashomon set of many near-optimal ones. We show t...
By Ethan Hsu, Harry Chen, Chudi Zhong, Lesia Semenova
The article discusses the necessity of building trust in AI for railway applications, highlighting that AI is currently limited to non‑safety critical uses due to stringent industry standards. It proposes focusing on three key areas—robustness, Operational Design Domain (ODD), and explainability—to meet compliance and safety requirements. By integrating these domains within a safe MLOps environment, the authors argue that regulatory acceptance and public confidence can be achieved, enabling broader AI adoption in mission‑critical railway systems.
By Lefebvre Renard Cl\'ement, L\'eb\'e Vincent, Da Silva Ribeiro Pereira Ricardo, Sundell Johan, Jaoul Arnaud Saiah Kenza, Mijatov\'ic Nenad
arXiv:2609.40034v1 Announce Type: cross
Abstract: Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM)...
By Ayoub Ajarra, Debabrota Basu
arXiv:2603. 24742v2 Announce Type: replace-cross Abstract: As the capabilities and adoption of Artificial Intelligence (AI) systems grow, trust in these AI systems is an increasingly urgent concern.
By Adeela Bashir, Zhao Song, Ndidi Bianca Ogbo, Nataliya Balabanova, Martin Smit, Chin-wing Leung, Paolo Bova, Manuel Chica Serrano, Dhanushka Dissanayake, Manh Hong Duong, Elias Fernandez Domingos, Nikita Huber-Kralj, Marcus Krellner, Andrew Powell, Stefan Sarkadi, Fernando P. Santos, Zia Ush Shamszaman, Chaimaa Tarzi, Paolo Turrini, Grace Ibukunoluwa Ufeoshi, Victor A. Vargas-Perez, Alessandro Di Stefano, Simon T. Powers, The Anh Han
arXiv:2607. 14152v1 Announce Type: cross Abstract: The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models.
By Michael Correll, Lucy Havens, Mahsan Nourani
arXiv:2607. 14315v1 Announce Type: cross Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score.
By Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, Jonh Soldatos
The article discusses how artificial intelligence is reshaping measurement in economics by converting unstructured data into structured variables at low cost, enabling large‑scale measurement that was previously infeasible. It outlines three stages—discovery, construct definition, and observation—where AI impacts the measurement pipeline and stresses the importance of rigorous validation to ensure credible inference. The review offers guidance on navigating the shift from a single scalable measure to multiple plausible ones that can lead to differing empirical conclusions.
By Melissa Dell, Ashesh Rambachan