arXiv AI
3d ago

Fine-Grained Multi Image Object Hallucination Benchmark

The paper introduces MIOH, a fine‑grained benchmark for evaluating object hallucination in multimodal large language models (MLLMs) across multiple images. It assesses hallucination through four tasks—existence, counting, attribute, position—using three reasoning patterns (comprehensive, comparative, selective) and three adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Evaluation of 29 models, including GPT‑5 and Gemini‑2.5‑Pro, shows distinct failure patterns, indicating that hallucination arises from integration‑stage limitations rather than just perceptual errors.

By Joonki Min, Chaeyun Kim, Hyungwook Choi, Yejin Kim, Kihyun Kim, Yohan Jo, Joonseok Lee
arXiv AI
Aug 12

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605. 16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu
arXiv AI
Aug 7

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

arXiv:2605. 16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu