arXiv AI

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

The study shows that hallucination detection in large language models is largely driven by a single mean‑shift component in hidden states. Across three 7B‑scale models and multiple datasets, removing this direction reduces detection to chance, while a simple L2‑regularized logistic regression achieves high AUROC (0.952) and outperforms more complex probe architectures. The authors introduce LayerMix, a multi‑layer aggregation method that matches oracle‑layer performance without requiring oracle access, demonstrating that apparent probe complexity stems from high‑dimensional covariance estimation rather than non‑linearity.

arXiv Machine Learning
6d ago

Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models

The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.

By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
arXiv AI
Aug 5

UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.

By Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
arXiv Computation and Language
Sep 10

Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection

The paper describes a system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. It combines a small 4‑B‑parameter VLM fine‑tuned as a per‑token classifier that uses a two‑token feature from its hidden states with a large ~400‑B zero‑shot VLM judge at prediction time, both leveraging OCR of visible in‑image text. Using synthetic hallucination data from the large model for ensemble diversity and validation‑based selection of feature layer, training data, and OCR grounding, the entry achieved competitive results across multiple languages.

By Eli Schwartz
arXiv AI
Aug 20

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

The paper introduces a token‑level hallucination detector that treats hallucinations as temporally extended spans and uses sequence labeling. It fuses 33‑dimensional features from text statistics, NLI entailment, and language‑model surprisal, and applies a BiGRU to achieve an AUC of 0.840 on RAGTruth, outperforming a logistic‑regression baseline by 11 points. The study shows that temporal ordering of features, rather than model capacity, drives most of the performance gain, and the detector remains effective on unseen language models with less than 4% AUC loss.

By Igor Itkin
arXiv Machine Learning
Aug 12

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

arXiv:2608. 10835v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input.

By Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca, Ethan Fetaya, Yftah Ziser, Gal Chechik, Haggai Maron
arXiv AI
6d ago

Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

The paper investigates the existence of hallucination neurons in large language models by applying a rigorous five‑step diagnostic protocol to sparse probing methods. Using open‑source LLMs on TriviaQA, BioASQ, and NQ‑Open, the authors find that while detection of hallucination neurons replicates across models and datasets, the neurons are not uniquely localized—showing high feature correlation, moderate bootstrap stability, and weak overlap between sparse and dense rankings. The study demonstrates that sparse predictive structure can coexist with non‑unique neuron selection, underscoring the need for routine diagnostic validation in mechanistic interpretability.

By Huseyin Cavus, Sebin Sabu, Joshua Spear, Jaskaran Singh Kawatra, Pavithra Rajendran