RECAP: Relation Evidence Calibration for Detecting Spatial Relation Hallucinations in Vision-Language Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
ReWEIGH the Evidence is a training‑free decoding technique that calibrates token‑level ordinal visual evidence to reduce hallucinations in large vision‑language models. It aggregates vocabulary ranks across visual positions, compares candidates to a token‑specific reference derived from unlabeled images, and applies a bounded penalty only when evidence falls below this reference. Experiments on four 7B backbones show up to a 21.3% reduction in hallucinated object mentions while largely preserving or improving descriptive and general performance, with minimal added latency.
arXiv:2608. 03817v1 Announce Type: cross Abstract: Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence.
RelCheck is a training‑free post‑hoc correction pipeline that addresses relational hallucinations in multimodal large language models. It augments object‑level visual grounding with two forms of relational evidence—learned scene‑graph triples from RelTR and deterministic spatial predicates derived from bounding‑box geometry—forming a three‑layer visual knowledge base. When applied to LLaVA v1 13B, RelCheck improves the overall MME hallucination score from 585.0 to 630.0, with the most significant gain on spatial position accuracy.
arXiv:2607. 27069v2 Announce Type: cross Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts.
arXiv:2608.29193v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral...