arXiv Computer Vision By Jiale Song, Jiaxin Luo, Xue-song Tang, Kuangrong Hao, Mingbo Zhao

Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models

Read the original on arXiv Computer Vision →

The paper introduces Segmentation-Based Attention Entropy (SAE), a method that uses semantic segmentation to measure visual attention uncertainty in large vision‑language models (LVLMs). SAE provides a reliability score for detecting hallucinated objects and an attention‑adjustment technique that reduces hallucinations during inference. Experiments on public benchmarks and real quadruped robot scenarios demonstrate that SAE improves LVLM reliability without requiring additional training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 16

Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification

arXiv:2603. 24058v2 Announce Type: replace-cross Abstract: Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a critical barrier to their deployment in high-stakes scenarios such as autonomous driving and medical image analysis.

By Han Sun, Qin Li, Peixin Wang, Min Zhang