arXiv AI

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

arXiv:2605. 16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

arXiv AI
Aug 12

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605. 16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu
arXiv AI
Jul 16

Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

arXiv:2607. 13712v1 Announce Type: cross Abstract: Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility.

By Zhixiao Zheng, Zheren Fu, Zhiyuan Yao, Chunxiao Liu, Dongming Zhang, Zhendong Mao
arXiv Computer Vision
Sep 16

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.

By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
arXiv Machine Learning
Jun 26

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.

By Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, Tianyang Wang, Hao Xu
arXiv AI
Sep 15

Hallucination in Multimodal Foundation Models: A Survey on Causes, Corrections, and Evaluations

The article surveys hallucination issues in Large Vision‑Language Models (LVLMs), a type of multimodal foundation model that blends visual data with large language models. It categorizes hallucination causes into model architecture and data quality, presents a taxonomy of mitigation strategies, and critically evaluates existing evaluation benchmarks from both discriminative and generative viewpoints. The survey also outlines open challenges and future research directions to improve LVLM reliability and trustworthiness.

By Yinghao Guo, Wei Lan, Wenyi Chen, Qingfeng Chen, Shichao Zhang, Shirui Pan, Huiyu Zhou, Yi Pan