arXiv:2609.39142v1 Announce Type: new
Abstract: Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information th...
By Yunbei Zhang, Janet Wang, Jihun Hamm, Chandan K Reddy
arXiv:2609.39195v1 Announce Type: new
Abstract: Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own mo...
By Shichao Li, Meiqi Wang, Fei Su, Zhicheng Zhao
arXiv:2609.39266v1 Announce Type: new
Abstract: Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotat...
By Qixing Zhao, Jinpeng Li
arXiv:2609.39486v1 Announce Type: new
Abstract: Estimating accident mechanics from real-world crashes is important for vehicle-safety analysis, injury modeling, crash-severity prediction, and operati...
By Ond\v{r}ej Valach, V\'aclav Divi\v{s}, Ivan Gruber
arXiv:2609.39566v1 Announce Type: new
Abstract: Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected e...
By Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
arXiv:2609.39662v1 Announce Type: new
Abstract: Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity j...
By Eunmin Lee, Jungwoo Kim, Jong-Seok Lee
arXiv:2609.39794v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains rema...
By Zaijing Li, Rui Shao, Bing Hu, Haoyu Zhang, Dongmei Jiang, Liqiang Nie
arXiv:2609.39915v1 Announce Type: new
Abstract: Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents...
By Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie
arXiv:2609.39953v1 Announce Type: new
Abstract: Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computa...
By Jianghao Wang, Ke Meng, Jian Li, Chi Cheng, Longyu Qi, Liyin Liang, Yifeng Qian, Chunbo Lai, Yutian Lin, Zeyu Wang
arXiv:2609.40048v1 Announce Type: new
Abstract: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget,...
By Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
arXiv:2609.38494v1 Announce Type: cross
Abstract: Robotic manipulation integrates vision, touch, and language, whose importance shifts across stages: vision guides reaching, while touch, through its...
By Amir-Hossein Shahidzadeh, Seungjae Lee, Eadom Dessalene, Shanthosh Raaj Mohanram Mageswari, Soroush Etemad, Furong Huang, Cornelia Ferm\"uller, Yiannis Aloimonos
arXiv:2609.38616v1 Announce Type: cross
Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...
By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv:2609.39324v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. Howeve...
By Jingqiu Wang, Yan Wang
arXiv:2601.08355v3 Announce Type: replace
Abstract: Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for sa...
By Guo Cheng, Huang Li
arXiv:2601.21857v2 Announce Type: replace
Abstract: We revisit diffusion-based generation for structured visual content and identify a fundamental limitation of existing approaches: foreground preser...
By Taewon Kang, Yu Shen, Ming C. Lin
arXiv:2603.17343v3 Announce Type: replace
Abstract: The rapid proliferation of AI-Generated Images (AIGIs) poses severe misinformation risks, making AIGI detection critical yet challenging. Tradition...
By Chenyang Zhu, Maorong Wang, Jun Liu, Ching-Chun Chang, Isao Echizen
arXiv:2603.24528v2 Announce Type: replace
Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...
By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
arXiv:2604.18484v2 Announce Type: replace
Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from co...
By Kangan Qian, ChuChu Xie, Yang Zhong, Jingrui Pang, Siwen Jiao, Sicong Jiang, Zilin Huang, Yunlong Wang, Kun Jiang, Mengmeng Yang, Hao Ye, Guanghao Zhang, Hangjun Ye, Guang Chen, Long Chen, Diange Yang
arXiv:2605.07562v2 Announce Type: replace
Abstract: Remote sensing imagery spans ground sampling distances (GSDs) from centimeters to tens of meters, so both the visual evidence for a geographic conc...
By Song Zhang, Yanlong Chen, Yining Chen, Xiaowei Zhang, Yawei Li
arXiv:2605.22208v3 Announce Type: replace
Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly sele...
By Kailin Zhuang, Jiawei Wu, Zhi Jin