arXiv AI

Robust Explanations for User Trust in Enterprise NLP Systems

arXiv:2604. 12069v3 Announce Type: replace-cross Abstract: Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common case of black-box deployment (API-only access) where representation-based explainers are infeasible and existing studies provide limited guidance on whether explanations remain stable under real user noise, especially when organizations migrate from encoder classifiers to decoder LLMs.

arXiv AI
Aug 24

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

The paper introduces Trustworthy RAG, an evaluation agent designed to detect misinformation and knowledge poisoning in Retrieval-Augmented Generation systems. It combines natural language inference verification, a five-signal poison detector, and a weighted Trust Index to assess the reliability of retrieved content. Experiments on multiple LLMs show high accuracy and precision, with the agent effectively blocking unsafe advice in a secure-coding assistant scenario.

By Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
arXiv AI
Aug 11

Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation

arXiv:2510. 21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form responses generated by LLMs.

By Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner
arXiv AI
Aug 26

A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification

The paper introduces a formal auditing framework to evaluate the robustness and fidelity of post‑hoc explainers such as SHAP and LIME. It defines a Trust Score that combines how stable an explanation is under small input perturbations with how well the highlighted features actually influence the model’s prediction. Experiments on a Madagascar malnutrition dataset show that even highly accurate models can produce unreliable explanations, and that fidelity scores degrade when models overfit.

By Rosa Elysabeth Ralinirina, Jean Christian Ralaivao, Niaiko Micha\"el Ralaivao, Alain Josu\'e Ratovondrahona, Thomas Mahatody
arXiv AI
Sep 15

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

The paper introduces CarryOnBench, an interactive benchmark that tests whether large language models can revise their interpretation of user intent and recover utility while staying safe in multi‑turn conversations. Using 398 harmful‑looking queries with benign intents, the benchmark simulates 5,970 conversations across 14 models, evaluating both intent‑aligned utility and safety with a new metric called Ben‑Util. Results show that models often withhold information due to misinterpretation, but most can recover with clarifications, revealing failure modes such as unsafe and redundant recovery that single‑turn tests miss.

By Mingqian Zheng, Malia Morgan, Liwei Jiang, Carolyn Rose, Maarten Sap
arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo