arXiv:2606. 02628v1 Announce Type: new Abstract: We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal is strongest.
By Aizierjiang Aiersilan
The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.
By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.
By Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
The paper describes a system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. It combines a small 4‑B‑parameter VLM fine‑tuned as a per‑token classifier that uses a two‑token feature from its hidden states with a large ~400‑B zero‑shot VLM judge at prediction time, both leveraging OCR of visible in‑image text. Using synthetic hallucination data from the large model for ensemble diversity and validation‑based selection of feature layer, training data, and OCR grounding, the entry achieved competitive results across multiple languages.
By Eli Schwartz
The paper introduces a token‑level hallucination detector that treats hallucinations as temporally extended spans and uses sequence labeling. It fuses 33‑dimensional features from text statistics, NLI entailment, and language‑model surprisal, and applies a BiGRU to achieve an AUC of 0.840 on RAGTruth, outperforming a logistic‑regression baseline by 11 points. The study shows that temporal ordering of features, rather than model capacity, drives most of the performance gain, and the detector remains effective on unseen language models with less than 4% AUC loss.
By Igor Itkin
arXiv:2606. 24790v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinations.
By Anand Kamat, Daniel Blake, Brent M. Werness