Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.
arXiv:2608. 07302v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image.
The paper introduces MIOH, a fine‑grained benchmark for evaluating object hallucination in multimodal large language models (MLLMs) across multiple images. It assesses hallucination through four tasks—existence, counting, attribute, position—using three reasoning patterns (comprehensive, comparative, selective) and three adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Evaluation of 29 models, including GPT‑5 and Gemini‑2.5‑Pro, shows distinct failure patterns, indicating that hallucination arises from integration‑stage limitations rather than just perceptual errors.
arXiv:2606. 31054v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image.
arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.
arXiv:2607. 24017v1 Announce Type: cross Abstract: The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws.