arXiv:2606. 15753v1 Announce Type: new Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning.
By Yaoting Huang, Yifu Yuan, Linqi Han, Chengwen Li, Shuoheng Zhang, Xianze Yao, Hongyao Tang, Yan Zheng, Jianye Hao
arXiv:2507.00754v3 Announce Type: replace
Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...
By Selim Kuzucu, Muhammad Ferjad Naeem, Anna Kukleva, Federico Tombari, Bernt Schiele
arXiv:2602. 14065v2 Announce Type: replace Abstract: Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval.
By Kai Ye, Xianwei Mao, Sheng Zhou, Zirui Shao, Ye Mo, Liangliang Liu, Haikuan Huang, Bin Li, Jiajun Bu
WeakMCN introduces a multi-task collaborative network that jointly learns weakly supervised referring expression comprehension (WREC) and segmentation (WRES) using a dual-branch architecture. The WREC branch employs anchor-based contrastive learning and serves as a teacher for the WRES branch, while two novel modules—Dynamic Visual Feature Enhancement (DVFE) and Collaborative Consistency Module (CCM)—facilitate cross-task collaboration. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show significant performance gains over single-task baselines, with up to 3.91% and 13.11% improvements on WREC and WRES respectively, and strong generalization in semi-supervised settings.
By Silin Cheng, Yang Liu, Xinwei He, Sebastien Ourselin, Lei Tan, Gen Luo
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.
By Yanxiang Huang, Guohua Gao, Zhaoyang Wei
arXiv:2606. 30498v1 Announce Type: cross Abstract: Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color.
By Laines Schmalwasser, Jan Blunk, Niklas Penzel, Julia Niebling, Joachim Denzler
arXiv:2606. 19489v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space.
By Ya Wang, Adrian Paschke
arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.
By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv:2605.05593v2 Announce Type: replace
Abstract: Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the internal mechanisms governing how they encode...
By Zehao Deng, Tianjie Ju, Zheng Wu, Liangbo He, Jun Lan, Huijia Zhu, Weiqiang Wang, Zhuosheng Zhang
arXiv:2511. 18121v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information.
By Ming Zhong, Yuanlei Wang, Liuzhou Zhang, Ruichuan An, Renrui Zhang, Hao Liang, Ming Lu, Ying Shen, Wentao Zhang
arXiv:2603. 23867v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts.
By Weixin Chen, Antonio Vergari, Han Zhao
arXiv:2603. 07131v4 Announce Type: replace-cross Abstract: Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis.
By Shuai Lu, Meng Wang, Jia Guo, Jiawei Du, Bo Liu, Shengzhu Yang, Weihang Zhang, Huazhu Fu, Huiqi Li