arXiv Computer Vision
Sep 16

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.

By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
arXiv AI
Aug 12

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605. 16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu
arXiv Computation and Language
Sep 1

Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective

The paper argues that stochastic decoding, common in large language models, is not ideal for Visual Question Answering (VQA) because VQA is a closed‑ended task with head‑heavy answer distributions and epistemic uncertainty. The authors formalize how model calibration relates to predictive accuracy and identify conditions under which greedy decoding is optimal. Experiments across multiple benchmarks show greedy decoding outperforms stochastic sampling, and a new Greedy Decoding for Reasoning Models further improves multimodal reasoning performance.

By Boqi Chen, Xudong Liu, Yunke Ao, Jianing Qiu
arXiv AI
Aug 7

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

arXiv:2605. 16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu