arXiv Machine Learning

Virgil: Navigating Explainability for Transformer-based Language Models

Virgil is an interactive system designed to help practitioners and researchers navigate the growing but fragmented ecosystem of explainability tools for transformer-based language models. It provides a unified interface backed by a curated knowledge base, allowing users—including non-experts—to discover and compare different explainability tools. The tool aims to simplify access to these resources as transformer models are increasingly deployed in high‑stakes applications.

arXiv Machine Learning
Jun 2

Interpreto: An Explainability Library for Transformers

arXiv:2512. 09730v3 Announce Type: replace-cross Abstract: Interpreto is an open-source Python library for interpreting HuggingFace language models, from early BERT variants to LLMs.

By Antonin Poch\'e, Thomas Mullor, Gabriele Sarti, Fr\'ed\'eric Boisnard, Corentin Friedrich, Charlotte Claye, Fran\c{c}ois Hoofd, Raphael Bernas, Nicholas Asher, C\'eline Hudelot, Fanny Jourdan
arXiv Machine Learning
Sep 11

Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models

The paper introduces the Explainability Assistant, an open‑source conversational XAI system designed to interpret complex energy consumption forecasting models. By leveraging large language model function‑calling, it achieves 94% intent‑parsing accuracy and supports flexible natural‑language interaction across different ML problem types without task‑specific fine‑tuning. A comparative evaluation with energy domain specialists shows improved usability and consistent task accuracy, with all experts preferring the conversational interface over a traditional XAI dashboard.

By Rodion Krjut\v{s}kov, Eduard Barbu, Nikos Sakkas, Sofia Yfanti
arXiv AI
Jul 17

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

arXiv:2607. 14315v1 Announce Type: cross Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score.

By Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis, Dimitrios Kotios, Vasileios Koukos, Dimosthenis Kyriazis, Jonh Soldatos
arXiv Computation and Language
Sep 1

MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.

By Alireza Bayat Makou, Emirhan B\"oge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych
Hugging Face Trending Papers
Jun 9

Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions

As artificial intelligence and machine learning (AI/ML) models become integral to network operations, their lack of transparency poses a significant barrier to operator trust. Existing explainable artificial intelligence (XAI) techniques often fail to bridge this gap for non-specialists, producing technical outputs that are difficult to translate into actionable insights.