The paper introduces contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), a formalism used to represent and reason with information. Unlike traditional explanations that focus on a single argument, contrastive explanations highlight the differences between two topic arguments. The authors propose a general form of contrastive attribution functions (CAFs), present CAFs based on removal, gradients, and Shapley-values, and demonstrate their applicability in healthcare and bias identification contexts.
By Xiang Yin, Nico Potyka, Antonio Rago, Francesca Toni
arXiv:2605. 20098v2 Announce Type: replace Abstract: Claim verification is an important problem in high-stakes settings, including health and finance.
By Gabriel Freedman, Adam Dejl, Adam Gould, Mansi, Lihu Chen, Junqi Jiang, Francesca Toni
The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.
By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia
arXiv:2606. 16786v1 Announce Type: new Abstract: Algorithmic explanations are intended to help stakeholders understand opaque algorithmic decisions, but in practice, they often fall short.
By Eric G\"unther, Bal\'azs Szabados, Kristof Meding, Gunnar K\"onig, Sebastian Bordt, Ulrike von Luxburg
arXiv:2606. 09030v1 Announce Type: cross Abstract: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify.
By Hyeongwon Jang, Gyouk Chu, Changhun Kim, Joonhyung Park, Hangyul Yoon, Eunho Yang
The paper introduces a hypothesis‑testing framework that embeds feature importance methods (FIMs) within a Weight of Evidence (WoE) analysis. By quantifying how strongly observed evidence supports a given hypothesis—whether from domain knowledge, ground truth, or the FIM itself—the approach evaluates FIM alignment and variability. The authors provide theoretical links between WoE and attribution variance and demonstrate the method on LIME and SHAP explanations across varied reference hypotheses.
By Eddie Conti, Claudio Daka, \'Alvaro Parafita, Antonio L. Alfeo, Axel Brando, Mario G. C. A. Cimino
arXiv:2608. 07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each.
By Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, Francesca Toni
arXiv:2606. 06646v1 Announce Type: cross Abstract: Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics.
By Jakub B\k{a}ba, Jaros{\l}aw Chudziak
arXiv:2607. 09664v1 Announce Type: new Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentation.
By Anca Marginean, Adrian Groza
arXiv:2508.08966v2 Announce Type: replace
Abstract: The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a gro...
By Marte Eggen, Jacob Lysn{\ae}s-Larsen, Inga Str\"umke
arXiv:2504. 19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance.
By Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh