arXiv Machine Learning

Reliability, Faithfulness, and the Limits of Post-hoc Explanations of Opaque Scientific Models

arXiv:2606. 29346v1 Announce Type: new Abstract: Post-hoc explanation methods are routinely used to interpret scientific machine learning models, with the deliverable understood to be insight into the phenomenon the model has been trained on.

arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody
arXiv Computer Vision
Sep 7

From Interpretability Methods to Interpretable Models

The paper argues that explainable AI for computer vision has focused too much on developing interpretability methods rather than assessing how interpretable the models themselves are. It proposes a shift toward model-centric evaluation, using existing tools to compare what different models represent and compute, and emphasizes the need to measure whether humans can truly understand these models. The authors review the current toolbox, survey limited model comparison work, draw parallels to systems neuroscience, and outline a future agenda for model-focused XAI.

By Julien Colin, Nuria Oliver, Thomas Serre
arXiv AI
Sep 7

Evidence Integration in Large Language Models

The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.

By Sebastien Kawada, Manolis Kellis