arXiv:2606. 20077v1 Announce Type: cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals.
By Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito
arXiv:2602.06652v2 Announce Type: replace
Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...
By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv:2508.03351v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse language tasks, motivating their extension to vision-la...
By Yufei Xue, Yushi Huang, Lunjie Zhu, Jiawei Shao, Jun Zhang
OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.
By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
arXiv:2608.22916v1 Announce Type: new
Abstract: Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unc...
By Zeyu Wang, Xinming Xu
arXiv:2605. 20950v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) face a bottleneck of prohibitive computational costs arising from massive visual token sequences during inference.
By Yulin Zhao, Zheng Zhang