arXiv AI

GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS

GHOST-Q evaluates how post‑training quantization affects visual grounding in vision‑language models. The study compares three 8B VLM families across FP16, INT8, and NF4 precisions, pairing predictions to measure how compression redistributes grounding successes and failures. While most quantized variants maintain overall accuracy, several exhibit significant changes in hallucination‑sensitive conditions, and memory savings do not always translate to lower latency.

arXiv Computer Vision
Sep 3

Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

The paper examines six inference-time hallucination mitigation methods applied to three large vision-language models across four benchmarks, including MMStar. It finds that reducing hallucination rates often comes at the cost of lower informativeness—such as decreased object recall, visual coverage, and response detail—and that gains on hallucination benchmarks do not consistently translate to improved performance on fine-grained perception and reasoning tasks. The authors argue that current evaluation protocols may overstate progress by favoring conservative generation, and propose that hallucination mitigation should be assessed as a trade-off among faithfulness, informativeness, and overall capability.

By Mehrdad Fazli, Sina Mansouri, Mohit Marvania, Ziwei Zhu
arXiv Machine Learning
6d ago

Where Hallucinations Live: A Cross-Architecture Circuit in VQ-Tokenized Vision-Language Models

The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.

By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
arXiv Computer Vision
Aug 31

Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models

The paper introduces Dynamic Alignment Compensation (DAC), a training‑free inference‑time technique designed to reduce hallucinations in Large Vision‑Language Models (LVLMs). DAC monitors cross‑modal representation drift across decoder layers and generation steps, applying lightweight residual compensation through Layer‑wise Semantic Compensation and Sequential Semantic Correction. Experiments on nine multimodal benchmarks across various LVLM backbones demonstrate that DAC consistently lowers hallucination rates while preserving overall performance.

By Kairong Yu, Zixin Zhu, Le Yu, Hongwei Wang
arXiv Computer Vision
Sep 3

Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification

The paper introduces CADMP, a lightweight framework for detecting object hallucinations in large vision‑language models. CADMP measures cross‑modal attention drift between adjacent layers and verifies predictions by masking visually relevant regions, combining these signals to identify hallucinated outputs. Experiments on multiple benchmarks show that CADMP achieves competitive detection performance, and ablation studies confirm the complementary roles of attention drift and mask‑based verification.

By Xuanbing Wen, Boxu Chen, Le Yang, Jiakai Wang, Zhengyu Zhao, Chenhao Lin, Chao Shen
arXiv AI
Jun 16

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.

By Han Sun, Qin Li, Peixin Wang, Min Zhang
arXiv AI
Sep 15

Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations

The article surveys hallucination issues in Large Vision‑Language Models (LVLMs), a type of multimodal foundation model that blends visual data with large language models. It categorizes hallucination causes into model architecture and data quality, presents a taxonomy of mitigation strategies, and critically evaluates existing evaluation benchmarks from both discriminative and generative viewpoints. The survey also outlines open challenges and future research directions to improve LVLM reliability and trustworthiness.

By Yinghao Guo, Wei Lan, Wenyi Chen, Qingfeng Chen, Shichao Zhang, Shirui Pan, Huiyu Zhou, Yi Pan
arXiv AI
Aug 20

ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models

ReWEIGH the Evidence is a training‑free decoding technique that calibrates token‑level ordinal visual evidence to reduce hallucinations in large vision‑language models. It aggregates vocabulary ranks across visual positions, compares candidates to a token‑specific reference derived from unlabeled images, and applies a bounded penalty only when evidence falls below this reference. Experiments on four 7B backbones show up to a 21.3% reduction in hallucinated object mentions while largely preserving or improving descriptive and general performance, with minimal added latency.

By Jihae Jeong, Junha Choi, Hwanjo Yu