arXiv Computation and Language By Eli Schwartz

Two-Token Features and Small-Large Ensembles for VLM Hallucination Detection

Read the original on arXiv Computation and Language →

The paper describes a system for the SHROOM-Visions 2026 shared task on character-level VLM hallucination detection. It combines a small 4‑B‑parameter VLM fine‑tuned as a per‑token classifier that uses a two‑token feature from its hidden states with a large ~400‑B zero‑shot VLM judge at prediction time, both leveraging OCR of visible in‑image text. Using synthetic hallucination data from the large model for ensemble diversity and validation‑based selection of feature layer, training data, and OCR grounding, the entry achieved competitive results across multiple languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 1

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

SpanCalib-VLM is a hybrid system for detecting hallucinated text spans in Vision‑Language Models. It combines a multimodal sequence tagger (XLM‑RoBERTa‑Large + SigLIP) with a fine‑tuned generative VLM (Qwen3.5‑4B‑SHROOM‑SFT) and uses a Union‑Calibrated Fusion strategy to re‑score candidate spans. On the SHROOM‑Visions English evaluation split, the ensemble achieves a Pearson calibration correlation of 0.41, an overall IoU of 0.39, a clean‑response IoU of 0.91, and a detection accuracy of 70.7%.

By Amanuel Gizachew Abebe, Yasmin Moslem
arXiv Computation and Language
Aug 27

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

In 2026, the SHROOM-Visions shared task was launched at the UncertaiNLP Workshop co‑located with EMNLP to address hallucinations in large vision‑language models. The task builds on the SHEEP dataset and asks participants to detect and classify fine‑grained hallucination spans in image‑conditioned text generation across four languages (Chinese, English, French, Italian) using a five‑class taxonomy. The competition attracted 27 teams and over 600 system submissions, with top systems achieving character‑level, label‑conditioned, and IoU scores of 0.58, 0.46, and 0.51 respectively, surpassing baselines by 30‑40 points.

By Ra\'ul V\'azquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Cal\`o, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, J\"org Tiedemann, Timothee Mickus