arXiv AI

Explainability in Practice: A Survey of Explainable NLP Across Various Domains

arXiv:2502. 00837v3 Announce Type: replace-cross Abstract: Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management, where models such as GPT-4o, Gemini, and BERT increasingly inform decisions.

arXiv AI
2d ago

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

The paper introduces a comparative explainability framework for auditing DeBERTa‑v3 in zero‑shot medical abstract classification. It evaluates five explanation methods—SHAP, LIME, occlusion, Input × Gradient, and Attention × Gradient—using a natural language inference engine on a balanced corpus of 1,000 abstracts per diagnostic category. The study finds that explanatory stability aligns with predictive certainty, identifies three systemic failure mechanisms, and recommends combining multiple explanation methods and quantitative agreement metrics for transformer‑based medical text classifiers.

By Javier Diaz Esteban-Herreros, David Mu\~noz-Valero, Raquel Mart\'inez-Espa\~na, Jose M. Juarez, Juan Moreno-Garcia
arXiv AI
Sep 11

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

XAI-Arena proposes using large language models (LLMs) as judges to evaluate the quality of explainable AI (XAI) explanations, aiming for reproducibility, scalability, and multidimensional assessment. The framework assesses dimensions such as simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability across different datasets, models, and stakeholder personas. Human validation shows a strong positive correlation between LLM-generated and human ratings (Spearman's rho = .693, p < .001), supporting the viability of LLM-based evaluations.

By Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel
arXiv Machine Learning
Sep 11

Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models

The paper introduces the Explainability Assistant, an open‑source conversational XAI system designed to interpret complex energy consumption forecasting models. By leveraging large language model function‑calling, it achieves 94% intent‑parsing accuracy and supports flexible natural‑language interaction across different ML problem types without task‑specific fine‑tuning. A comparative evaluation with energy domain specialists shows improved usability and consistent task accuracy, with all experts preferring the conversational interface over a traditional XAI dashboard.

By Rodion Krjut\v{s}kov, Eduard Barbu, Nikos Sakkas, Sofia Yfanti
arXiv AI
Aug 19

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

The study investigates whether Large Language Models (LLMs) can translate technical explanations from credit risk models into stakeholder-friendly narratives. Using Freddie Mac loan data, the authors compare standard tabular models (XGBoost + SHAP) with alternative data pipelines (GNN + GNNExplainer and a bimodal mix) and generate explanations with three LLM configurations: a small fine‑tuned Gemma 3 4B, a large fine‑tuned DeepSeek R1 70B, and a zero‑shot Gemini 2.5. Findings show that the quality of explanations is more dependent on the evidence representation than on the LLM, that narratives reliably identify influential factors but are less consistent about the direction of influence, and that credit professionals demand higher evidentiary standards than non‑professionals.

By Sahab Zandi, Noah Kostesku, Christophe Mues, Mar\'ia \'Oskarsd\'ottir, Cristi\'an Bravo
arXiv AI
Jul 17

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

arXiv:2607. 14315v1 Announce Type: cross Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score.

By Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, Jonh Soldatos
arXiv AI
Jul 28

Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing

arXiv:2607. 23368v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful and interpretable explanations remains a key consideration for their responsible deployment in clinical settings.

By Jakub Rymarski (University of Warsaw, Poland), Adam Rempa{\l}a (University of Warsaw, Poland), Bart{\l}omiej Sobieski (University of Warsaw, Poland), Przemys{\l}aw Biecek (University of Warsaw, Poland)