The paper introduces contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), a formalism used to represent and reason with information, including augmenting AI classification tasks with explainability. Unlike traditional explanations that focus on a single argument, contrastive explanations highlight differences between two topic arguments. The authors propose general contrastive attribution functions (CAFs), present CAFs based on removal, gradients, and Shapley-values, analyze their properties, and demonstrate their applicability in healthcare and bias identification scenarios.
arXiv:2605. 20098v2 Announce Type: replace Abstract: Claim verification is an important problem in high-stakes settings, including health and finance.
By Gabriel Freedman, Adam Dejl, Adam Gould, Mansi, Lihu Chen, Junqi Jiang, Francesca Toni
arXiv:2606. 16786v1 Announce Type: new Abstract: Algorithmic explanations are intended to help stakeholders understand opaque algorithmic decisions, but in practice, they often fall short.
By Eric G\"unther, Bal\'azs Szabados, Kristof Meding, Gunnar K\"onig, Sebastian Bordt, Ulrike von Luxburg
The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.
By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia
The paper introduces a hypothesis‑testing framework that embeds feature importance methods (FIMs) within a Weight of Evidence (WoE) analysis. By quantifying how strongly observed evidence supports a given hypothesis—whether from domain knowledge, ground truth, or the FIM itself—the approach evaluates FIM alignment and variability. The authors provide theoretical links between WoE and attribution variance and demonstrate the method on LIME and SHAP explanations across varied reference hypotheses.
By Eddie Conti, Claudio Daka, \'Alvaro Parafita, Antonio L. Alfeo, Axel Brando, Mario G. C. A. Cimino
arXiv:2606. 06646v1 Announce Type: cross Abstract: Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics.
By Jakub B\k{a}ba, Jaros{\l}aw Chudziak
ProToMEx is a new explainability framework that uses Probabilistic Topic Models to learn latent topics representing high‑level reasons behind a classifier’s decisions, moving beyond simple feature attribution. It provides both global and local explanations, revealing multiple co‑existing reasons for individual predictions. Empirical results show that ProToMEx achieves comparable fidelity to SHAP and LIME while being 30–40× faster on standard tabular and synthetic datasets.
By Athina Georgara, Adarsh Valoor, Sarvapali D. Ramchurn
The paper investigates how dense embedding models can be used for stance-aware argument retrieval, a task that requires both topic relevance and correct stance (support or attack) toward a claim. Experiments reveal that current models favor topical overlap and ignore stance, and that contrastive training to fix this bias leads to over-correction, where models focus too much on polarity keywords at the expense of topic relevance. To address this, the authors propose diagnostic word-ablation metrics and a data‑centric solution involving a balanced argument curriculum and LLM‑augmented stance‑inverted arguments, which helps powerful models learn deeper directional logic and improves stance‑aware retrieval performance.
By Angelo Sparacino, Francesca Toni, Adam Dejl
arXiv:2508.08966v2 Announce Type: replace
Abstract: The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a gro...
By Marte Eggen, Jacob Lysn{\ae}s-Larsen, Inga Str\"umke
arXiv:2608.29529v1 Announce Type: cross
Abstract: Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstrac...
By William Schroeder
arXiv:2608. 25897v1 Announce Type: new Abstract: Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior.
By Xu Zheng, Zichuan Liu, Zhuomin Chen, Mayur Akewar, Janki Bhimani, Jason Liu, Mo Sha, Jingchao Ni, Wei Cheng, Dongsheng Luo
arXiv:2608. 07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each.
By Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, Francesca Toni