arXiv:2606. 07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.
By Lujun Li, Lama Sleem, Niccolo Gentile, Yangjie Xu, Yewei Song, Wenbo Wu, Radu State
arXiv:2605. 30170v2 Announce Type: replace-cross Abstract: While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting.
By Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
arXiv:2606. 20077v1 Announce Type: cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals.
By Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito
arXiv:2608. 10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects.
By Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
arXiv:2606. 22437v2 Announce Type: replace-cross Abstract: We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results.
By Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong, Chengping Zhao, Ting Liu, Yuzhuo Fu
arXiv:2607. 06420v1 Announce Type: cross Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning.
By Jinhong Deng, Limeng Qiao, Guanglu Wan