arXiv Computation and Language

Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection

The paper describes a system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. It combines a small 4‑B‑parameter VLM fine‑tuned as a per‑token classifier that uses a two‑token feature from its hidden states with a large ~400‑B zero‑shot VLM judge at prediction time, both leveraging OCR of visible in‑image text. Using synthetic hallucination data from the large model for ensemble diversity and validation‑based selection of feature layer, training data, and OCR grounding, the entry achieved competitive results across multiple languages.

arXiv Computation and Language
Sep 1

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

SpanCalib-VLM is a hybrid system for detecting hallucinated text spans in Vision‑Language Models. It combines a multimodal sequence tagger (XLM‑RoBERTa‑Large + SigLIP) with a fine‑tuned generative VLM (Qwen3.5‑4B‑SHROOM‑SFT) and uses a Union‑Calibrated Fusion strategy to re‑score candidate spans. On the SHROOM‑Visions English evaluation split, the ensemble achieves a Pearson calibration correlation of 0.41, an overall IoU of 0.39, a clean‑response IoU of 0.91, and a detection accuracy of 70.7%.

By Amanuel Gizachew Abebe, Yasmin Moslem
arXiv Computation and Language
Aug 27

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

In 2026, the SHROOM-Visions shared task was launched at the UncertaiNLP Workshop co‑located with EMNLP to address hallucinations in large vision‑language models. The task builds on the SHEEP dataset and asks participants to detect and classify fine‑grained hallucination spans in image‑conditioned text generation across four languages (Chinese, English, French, Italian) using a five‑class taxonomy. The competition attracted 27 teams and over 600 system submissions, with top systems achieving character‑level, label‑conditioned, and IoU scores of 0.58, 0.46, and 0.51 respectively, surpassing baselines by 30‑40 points.

By Ra\'ul V\'azquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Cal\`o, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, J\"org Tiedemann, Timothee Mickus
arXiv AI
Aug 5

UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space

arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.

By Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban
arXiv AI
Jul 7

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

arXiv:2607. 04163v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering.

By Kai Tang, Jinhao You, Bohua Zhang, Yichen Guo, Yiding Sun, Dongxu Zhang, Chenxi Li, Xiande Huang, Shanghang Zhang
arXiv Machine Learning
Aug 12

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

arXiv:2608. 10835v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input.

By Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca, Ethan Fetaya, Yftah Ziser, Gal Chechik, Haggai Maron
arXiv AI
Aug 20

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

The paper introduces a token‑level hallucination detector that treats hallucinations as temporally extended spans and uses sequence labeling. It fuses 33‑dimensional features from text statistics, NLI entailment, and language‑model surprisal, and applies a BiGRU to achieve an AUC of 0.840 on RAGTruth, outperforming a logistic‑regression baseline by 11 points. The study shows that temporal ordering of features, rather than model capacity, drives most of the performance gain, and the detector remains effective on unseen language models with less than 4% AUC loss.

By Igor Itkin
arXiv AI
Sep 1

The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

The study shows that hallucination detection in large language models is largely driven by a single mean‑shift component in hidden states. Across three 7B‑scale models and multiple datasets, removing this direction reduces detection to chance, while a simple L2‑regularized logistic regression achieves high AUROC (0.952) and outperforms more complex probe architectures. The authors introduce LayerMix, a multi‑layer aggregation method that matches oracle‑layer performance without requiring oracle access, demonstrating that apparent probe complexity stems from high‑dimensional covariance estimation rather than non‑linearity.

By Jungseob Lee, Jaehyung Seo, Heuiseok Lim