arXiv Machine Learning

Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction

The paper introduces a low-cost method for detecting hallucinations in large language models by treating the model as a black-box dynamical system. It projects responses into a high-dimensional manifold, models the latent state-space dynamics with Koopman operator theory, and uses differential residual scores from transition operators to distinguish factual from hallucinated outputs. The approach requires only a single-sample pass and shows state-of-the-art performance across three benchmarks with reduced resource overhead.

arXiv Computation and Language
Aug 25

DynHD: Hallucination Detection for Diffusion Large Language Models via Denoising Dynamics Deviation Learning

DynHD is a method for detecting hallucinations in diffusion large language models (D‑LLMs) by focusing on token‑level uncertainty and its evolution during the denoising process. It introduces a semantic‑aware evidence construction module that filters out non‑informative structural tokens and highlights uncertainty in informative tokens, and a reference evidence generator that models the expected trajectory of uncertainty, enabling a deviation‑based detector to identify hallucinations. Experiments show DynHD outperforms existing baselines while being more efficient across various benchmarks and backbone models.

By Yanyu Qian, Yue Tan, Yixin Liu, Wang Yu, Shirui Pan
arXiv AI
2d ago

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

The paper proposes a hidden‑state probing method for detecting hallucinations at the span level in large language model outputs, moving beyond token‑wise binary classification. By examining layer‑wise activation patterns, the approach identifies the exact onset and continuation tokens of hallucinations, achieving higher precision‑recall AUC than random baselines despite class imbalance. Additionally, the authors introduce a cross‑model detection framework where one model observes another’s internal representations, showing that an external observer can match or surpass the generator’s own self‑detection of hallucination onsets, even when the observer is smaller.

By Kingshuk Gupta, Davide Buscaldi
Hugging Face Trending Papers
Aug 20

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

The paper extends a dynamical systems approach to classify unsafe outputs from large language models (LLMs) by projecting prompts and responses into high‑dimensional embeddings and fitting separate Koopman-based predictive models for safe and unsafe regimes. A differential residual score compares prediction errors from these models to classify new outputs. Experiments on three safety benchmarks show that including prompt embeddings improves detection of interaction‑dependent violations, especially with causal decoders like Llama‑3, while response‑only violations benefit more from dense semantic embeddings.

arXiv Computation and Language
Aug 28

Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models

The paper introduces Prediction of Prediction (PoP), a method that fuses intermediate hidden representations across transformer layers during a single forward pass to detect hallucinations in large language models. PoP leverages internal hidden‑state transition dynamics to signal factual errors without extra decoding steps, achieving a 75.5% AUROC on the TruthfulQA benchmark with less than 1.2% added latency.

By Himal Badu
arXiv Computer Vision
Aug 31

Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

The paper introduces Dynamic Alignment Compensation (DAC), a training‑free inference‑time technique designed to reduce hallucinations in Large Vision‑Language Models (LVLMs). DAC monitors cross‑modal representation drift across decoder layers and generation steps, applying lightweight residual compensation through Layer‑wise Semantic Compensation and Sequential Semantic Correction. Experiments on nine multimodal benchmarks across various LVLM backbones demonstrate that DAC consistently lowers hallucination rates while preserving overall performance.

By Kairong Yu, Zixin Zhu, Le Yu, Hongwei Wang