arXiv AI By Yiyang Chen, Yixin Tan, Binrui Shen

Listening makes Vision Clear for VLMs

Read the original on arXiv AI →

arXiv:2606. 23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv Computer Vision
6d ago

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati