Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 07507v1 Announce Type: cross Abstract: Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence.
RelCheck is a training‑free post‑hoc correction pipeline that addresses relational hallucinations in multimodal large language models. It augments object‑level visual grounding with two forms of relational evidence—learned scene‑graph triples from RelTR and deterministic spatial predicates derived from bounding‑box geometry—forming a three‑layer visual knowledge base. When applied to LLaVA v1 13B, RelCheck improves the overall MME hallucination score from 585.0 to 630.0, with the most significant gain on spatial position accuracy.
arXiv:2603. 21693v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to hallucinations, defined as generating responses that contradict the input image, posing serious risks in clinical settings.
Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can still produce overconfident predictions and hallucination-like outputs, particularly when the visual evidence is weak, ambiguous, or semantically inconsistent.
arXiv:2608.21819v1 Announce Type: cross Abstract: Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions wh...
arXiv:2601.06847v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We po...