arXiv:2608. 14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation.
By Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
arXiv:2609.16646v1 Announce Type: new
Abstract: When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instrument...
By Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang
The paper introduces the Graded Color Attribution (GCA) dataset, a benchmark that tests whether Vision‑Language Models (VLMs) and humans can articulate and follow a threshold rule for labeling objects by color. In experiments, humans consistently adhere to their stated rules, while VLMs—despite accurately estimating color coverage—often violate their own introspective rules, especially when world‑knowledge priors are present. This discrepancy highlights a miscalibration in VLM self‑knowledge that differs from human cognition.
By Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
The article surveys hallucination issues in Large Vision‑Language Models (LVLMs), a type of multimodal foundation model that blends visual data with large language models. It categorizes hallucination causes into model architecture and data quality, presents a taxonomy of mitigation strategies, and critically evaluates existing evaluation benchmarks from both discriminative and generative viewpoints. The survey also outlines open challenges and future research directions to improve LVLM reliability and trustworthiness.
By Yinghao Guo, Wei Lan, Wenyi Chen, Qingfeng Chen, Shichao Zhang, Shirui Pan, Huiyu Zhou, Yi Pan
arXiv:2605.10893v3 Announce Type: replace
Abstract: Large vision-language models (LVLMs) suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven enti...
By Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Kundan Thind, Mohammad M. Ghassemi
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration...
HALDETECT is a system developed for the English hallucination-detection track of ImageEval 2026, where the task is to identify the single visually grounded statement among three culturally plausible options. The approach treats the problem as a contrastive decision, outputs the answer before an explanation, and bases reasoning on colour/texture, shape/form, and context. The best model fine‑tunes Qwen2.5‑VL‑7B‑Instruct with 4‑bit QLoRA, freezes the vision encoder, and achieves a Contrastive Instability score of 0.035 on the test set, placing third among eight teams.
By Syed Mohaiminul Hoque, Md Sakhawat Hossain
In 2026, the SHROOM-Visions shared task was launched at the UncertaiNLP Workshop co‑located with EMNLP to address hallucinations in large vision‑language models. The task builds on the SHEEP dataset and asks participants to detect and classify fine‑grained hallucination spans in image‑conditioned text generation across four languages (Chinese, English, French, Italian) using a five‑class taxonomy. The competition attracted 27 teams and over 600 system submissions, with top systems achieving character‑level, label‑conditioned, and IoU scores of 0.58, 0.46, and 0.51 respectively, surpassing baselines by 30‑40 points.
By Ra\'ul V\'azquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Cal\`o, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, J\"org Tiedemann, Timothee Mickus
arXiv:2504. 10020v4 Announce Type: replace-cross Abstract: Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs).
By Hao Yin, Guangzong Si, Zilei Wang
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding and complex reasoning. However, LVLMs do not just inherit the text-level hallucinations; they also hallucinate against the image, producing fluent responses ungrounded in what they see.
The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.
By Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza, Oleksandr Pryymak, Aryo Pradipta Gema, Iacopo Masi, Pasquale Minervini, Fabrizio Silvestri
arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.
By Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia, Mohamad Amin Yousefi, Asal Mohammadjafari Mamaqani, Behrad Samimi, Mohammad Hossein Rohban