arXiv AI

BEACON: Behavioral Entropy Aggregation for Cross-Model Hallucination Detection in Large Language Models

arXiv:2606. 07528v1 Announce Type: cross Abstract: Hallucination in large language models (LLMs), defined as the generation of factually incorrect or unsupported content, remains a critical barrier to reliable deployment.

arXiv AI
2d ago

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

The paper proposes a hidden‑state probing method for detecting hallucinations at the span level in large language model outputs, moving beyond token‑wise binary classification. By examining layer‑wise activation patterns, the approach identifies the exact onset and continuation tokens of hallucinations, achieving higher precision‑recall AUC than random baselines despite class imbalance. Additionally, the authors introduce a cross‑model detection framework where one model observes another’s internal representations, showing that an external observer can match or surpass the generator’s own self‑detection of hallucination onsets, even when the observer is smaller.

By Kingshuk Gupta, Davide Buscaldi
arXiv AI
Aug 20

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

The paper introduces a token‑level hallucination detector that treats hallucinations as temporally extended spans and uses sequence labeling. It fuses 33‑dimensional features from text statistics, NLI entailment, and language‑model surprisal, and applies a BiGRU to achieve an AUC of 0.840 on RAGTruth, outperforming a logistic‑regression baseline by 11 points. The study shows that temporal ordering of features, rather than model capacity, drives most of the performance gain, and the detector remains effective on unseen language models with less than 4% AUC loss.

By Igor Itkin
arXiv AI
Sep 3

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

The paper investigates hallucination detection in black‑box large language models by leveraging two accessible signals: semantic entropy, which captures disagreement among sampled response meanings, and token‑level uncertainty derived from log‑probabilities. It introduces a TopK aggregation technique, a hybrid CoCoA method combining uncertainty with semantic dissimilarity, and two supervised approaches—Gated and Stacked—that integrate token and semantic features. Across seven benchmarks and four language models, the supervised Stacked method performs best in many cases, while TopK and CoCoA remain competitive without labeled data, though all methods require careful threshold calibration.

By Urja Pawar, Rajitha Ramanayake, Owen O'Neill, Nabeel Kemal, Abhishek Mandal, Houssem Chatbri, Christopher Martin
Hugging Face Trending Papers
Aug 18

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

The paper introduces InnerExpert, a method that uses Mixture-of-Experts (MoE) internal signals—such as router entropy, expert disagreement, and usage patterns—to detect hallucinations at the token level in large language models. By combining these MoE-specific signals with standard transformer features into compact per-token vectors, InnerExpert trains a lightweight detector using an LLM-as-a-judge pipeline, enabling continuous updates without manual labeling. Experiments across five datasets and two MoE architectures show that InnerExpert outperforms existing methods, achieving up to 0.91 answer-level and 0.76 token-level AUROC with only a single forward pass.

arXiv AI
Aug 19

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

The paper introduces InnerExpert, a method that uses Mixture-of-Experts (MoE) architecture signals—such as router entropy, expert disagreement, and usage patterns—to detect hallucinations at the token level in Large Language Models. By combining these MoE-specific signals with standard transformer features into compact per-token vectors, InnerExpert trains a lightweight detector using an LLM-as-a-judge pipeline, enabling continuous updates without manual labeling. Experiments across five datasets and two MoE architectures show that InnerExpert outperforms existing methods, achieving up to 0.91 answer-level and 0.76 token-level AUROC with only a single forward pass.

By Joao Fonseca, Rodrigo Rodrigues, Paolo Romano