arXiv:2609.34863v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Thei...
By Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang, Jiayin Cai, Yao Hu
arXiv:2609.34220v2 Announce Type: replace-cross
Abstract: Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as ob...
By Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang
arXiv:2609.38855v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic m...
By Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng
arXiv:2609.39450v1 Announce Type: cross
Abstract: LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. Howe...
By Jihun Han, Yejin Jang, Byung Il Kwak, Mee Lan Han
arXiv:2609.40055v1 Announce Type: cross
Abstract: On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for...
By Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djuri\'c, Sima Mofakham
arXiv:2609.39665v1 Announce Type: new
Abstract: Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This...
By Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
arXiv:2609.38810v1 Announce Type: cross
Abstract: As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual sig...
By Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin
arXiv:2605.07872v2 Announce Type: replace-cross
Abstract: Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains s...
By Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang, Hao Zhou, Fandong Meng, Xu Sun
arXiv:2605.26380v2 Announce Type: replace-cross
Abstract: Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. How...
By Jingru Chen, Yiming Liu, Mingtao Chen, Sijie Chen, Richeng Xuan, Liang Yang, Zhichao Hu, Fanyang Lu
arXiv:2609.38716v1 Announce Type: cross
Abstract: Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weaknes...
By Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, Dongxiao Zhu
arXiv:2609.39178v1 Announce Type: cross
Abstract: Recently, Vision-Language-Action (VLA) models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understa...
By Songhua Yang, Ziyu Liu, Yuanwei Liu, Xuetao Li, Xuanye Fei, He Huang, Zheng Wang, Miao Li
arXiv:2609.38282v1 Announce Type: new
Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Seq...
By Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu
arXiv:2609.39168v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing...
By Zhihan Zhang, Lizi Liao
arXiv:2609.38465v1 Announce Type: cross
Abstract: Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that...
By Shuyang Jiang, Fucheng Deng, Yuchuan Luo, Zhenyu Wu
arXiv:2609.38899v1 Announce Type: cross
Abstract: Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit poli...
By Wenyu Chen, Li Wang, Chuanchao Zang, Xiangtao Meng, Xinyu Gao, Jianing Wang, Zheng Li, Shanqing Guo
arXiv:2609.39306v1 Announce Type: cross
Abstract: Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our...
By Shengjie Jin, Hengbo Xu, Zelong Sun, YuJie Guo, Zhiwu Lu
arXiv:2609.39924v1 Announce Type: cross
Abstract: Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, m...
By Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, Dianhai Yu
arXiv:2609.40079v1 Announce Type: cross
Abstract: While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to...
By Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao, Yitong Liu, Jiyin Li, Ruiqi Tang, Zixuan Lin, Yi Luo, Xurui Zhang, Ronghao Chen, Huacan Wang, Lei Li
arXiv:2609.40195v1 Announce Type: cross
Abstract: Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spann...
By Guangzhi Xiong, Xinyuan Zhang, Xiao Yang, Hyokun Yun, Kai Zhang, Shiun-Zu Kuo, Hyeonjeong Ha, Xilun Chen, Kai Sun, Lucas Liang, Guangqiang Dong, Ejaz Ahmed, Ahmed A Aly, Anuj Kumar, Raffay Hamid, Aidong Zhang, Xin Luna Dong
arXiv:2511.22232v2 Announce Type: replace-cross
Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical inter...
By Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen