arXiv:2609.24470v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly explored as interfaces for scientific image analysis, where a visual question-answering (VQA)...
By Samia Mohinta, Albert Cardona
arXiv:2608. 06270v1 Announce Type: new Abstract: The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom.
By Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu
arXiv:2609.22588v1 Announce Type: new
Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they h...
By Yuyang Dai, Bofei Huang, Hongbo Zhang, Haoran Xie
arXiv:2609.40325v1 Announce Type: new
Abstract: As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anoma...
By Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang
arXiv:2609.00232v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmar...
By Yue Zhou, Yuan Wu, Yi Chang
arXiv:2607. 15205v1 Announce Type: cross Abstract: Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task.
By Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng