arXiv Machine Learning

Interpreto: An Explainability Library for Transformers

arXiv:2512. 09730v3 Announce Type: replace-cross Abstract: Interpreto is an open-source Python library for interpreting HuggingFace language models, from early BERT variants to LLMs.

arXiv Computation and Language
Sep 7

NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution

NOTAI.AI is an explainable AI-generated text detection system that goes beyond a simple binary label by showing which signals influenced its prediction. It combines sentence-level conditional probability curvature, a neural detector score, and interpretable stylometric and readability features in an XGBoost meta-classifier, and explains predictions using TreeSHAP feature contributions that can be turned into concise natural-language explanations. Evaluated on a balanced subset of RAID, the full model achieves 0.9685 F1 and receives 94.5–98.6% approval from model judges for the faithfulness of its explanations.

By Oleksandr Marchenko Breneur, Adelaide Danilov, Aria Nourbakhsh, Salima Lamsiyah
arXiv Computation and Language
Sep 1

MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

MURANO is an open‑source framework that enables researchers to design, run, and reproduce mechanistic interpretability experiments on large language models. It unifies the five key stages—loading, recording, attribution, intervention, and evaluation—into composable pipeline steps that exchange named artifacts and use canonical addresses for interoperability. The authors demonstrate the framework with reproductions of existing studies and a sparse autoencoder case study, showing its practical applicability across disciplines.

By Alireza Bayat Makou, Emirhan B\"oge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych
arXiv AI
6d ago

A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code

The paper presents a taxonomy-driven framework for identifying, categorizing, and explaining bias in AI-generated Python code. By extending an existing dataset and manually annotating bias categories and justifications, the authors evaluate both proprietary and open-source large language models (LLMs) for automated bias detection and explanation. Results show that models such as Gemini and Qwen3-coder achieve high classification accuracy and produce justification and code identification similarities that closely match human-authored reasoning.

By Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
arXiv Machine Learning
Aug 27

Virgil: Navigating Explainability for Transformer-based Language Models

Virgil is an interactive system designed to help practitioners and researchers navigate the growing but fragmented ecosystem of explainability tools for transformer-based language models. It provides a unified interface backed by a curated knowledge base, allowing users—including non-experts—to discover and compare different explainability tools. The tool aims to simplify access to these resources as transformer models are increasingly deployed in high‑stakes applications.

By Martino Ciaperoni, Sezer Kutluk, Benedetta Muscato, Marta Marchiori Manerba, Fosca Giannotti