arXiv:2605.19410v2 Announce Type: replace
Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-ho...
By Zilin Wang, Stella X. Yu
arXiv:2605.25333v3 Announce Type: replace
Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption...
By Tianshuo Xu, Yichen Xie, Depu Meng, Chensheng Peng, Quentin Herau, Bo Jiang, Yihan Hu, Wei Zhan
arXiv:2606.20092v3 Announce Type: replace
Abstract: Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-...
By Ganlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, Wengang Zhou, Yao Mu, Tai Wang
arXiv:2609.13006v2 Announce Type: replace
Abstract: Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take wi...
By Minh-Loi Nguyen, Xuan-Vu Le, Trung-Nghia Le, Tam V. Nguyen, Minh-Triet Tran, Thanh-Toan Do
arXiv:2609.33399v2 Announce Type: replace
Abstract: In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and...
By Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin, Xi Yu, Xusen Hei, DingBa Fu, Jiayuan Xie, Yi Cai
arXiv:2609.33462v2 Announce Type: replace
Abstract: Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use...
By Shriram Damodaran, Soumyaratna Debnath, Cheston Tan, Lin Wang
arXiv:2609.34581v2 Announce Type: replace
Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, w...
By Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
arXiv:2609.36913v1 Announce Type: cross
Abstract: Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propos...
By Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn
arXiv:2609.36277v1 Announce Type: new
Abstract: Existing protein geometry models typically represent molecular surfaces using local geometric features such as sampled points, normals, and curvature....
By Siyuan Chen, Cai Zhou, Jinrui Zhang, Zhaokang Liang, Taku Komura, Wojciech Matusik, Stephen Bates, Tommi Jaakkola, Wengong Jin, Peter Yichen Chen, Minghao Guo
arXiv:2609.36118v1 Announce Type: new
Abstract: Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone la...
By Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames, Yisong Yue
arXiv:2609.36576v1 Announce Type: new
Abstract: Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language...
By Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson
arXiv:2609.36572v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further e...
By Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang
arXiv:2609.36416v1 Announce Type: cross
Abstract: Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions....
By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Jackson Lee, Thomas Wolf, Pragna Mannam
arXiv:2609.36557v1 Announce Type: cross
Abstract: Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reaso...
By Janet Wang, Yunbei Zhang, Xiao Wang, Jihun Hamm
arXiv:2609.36651v1 Announce Type: cross
Abstract: Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length b...
By FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu, Bo Xu, Guoqi Li
arXiv:2609.36759v1 Announce Type: cross
Abstract: Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners...
By Chiyuan He, Zihuan Qiu, Fanman Meng, Chao Wang, Liangjiang Chen, Linfeng Xu, Qingbo Wu, Hongliang Li
arXiv:2609.37002v1 Announce Type: cross
Abstract: High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to a...
By Xijia Tao, Yihua Teng, Xinyu Fu, Cheng Gong, Ziru Liu, Xudong Xie, Rui Liu, Lingpeng Kong
arXiv:2609.37181v1 Announce Type: cross
Abstract: Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has e...
By Jin Chen, Yiming Jiang, Chongyang Xu, Modi Shi, Shijia Peng, Li Chen, Tianyu Li, Mu Xu, Yilun Chen, Steven Hoi, Hongyang Li
arXiv:2609.37349v1 Announce Type: cross
Abstract: Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate s...
By Yalun Wu, Bingzhou Wang, Boyang Wang, Peiying Wang, Shaojie He, Yunhan Wang, Shaozu Yuan, Jiawei Wang
arXiv:2609.37426v1 Announce Type: cross
Abstract: Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture....
By Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matari\'c