arXiv:2609.37519v1 Announce Type: cross
Abstract: Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched...
By Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh
arXiv:2609.37591v1 Announce Type: cross
Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time...
By Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang, Ziju Han, Yahong Han
arXiv:2609.37783v1 Announce Type: cross
Abstract: Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not...
By Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jacquelyn Burkell, Yuntian Deng, Karen Eltis, Maura R. Grossman, Vered Shwartz, Ebrahim Bagheri
arXiv:2609.37871v1 Announce Type: cross
Abstract: Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDr...
By Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou, Lisheng Wu, Xuewei Li, Zequn Qin, Xi Li
arXiv:2609.37938v1 Announce Type: cross
Abstract: Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggr...
By Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng
arXiv:2609.38028v1 Announce Type: cross
Abstract: Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple st...
By Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer
arXiv:2609.38178v1 Announce Type: cross
Abstract: Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this i...
By Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao
arXiv:2602.11678v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: struct...
By Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu, Xia Hua, Si Shi, F. Richard Yu
arXiv:2506.11261v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and...
By Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid
arXiv:2609.32292v2 Announce Type: replace-cross
Abstract: Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into relia...
By Xuekang Yang, Lu Chen, Shuang Luo, Jialing Zhu, Qi Zhang, Yue Gao, Xiang Zhang
arXiv:2609.33299v2 Announce Type: replace-cross
Abstract: World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the...
By Cunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou, Yue Yao, Chi Harold Liu
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
arXiv:2609.37432v1 Announce Type: new
Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also sup...
By T. Konstantin Rusch, Tim Seyde, Jared Boyer, Zach J. Patterson, Daniela Rus
arXiv:2609.38133v1 Announce Type: new
Abstract: Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general...
By Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti
arXiv:2609.36012v1 Announce Type: cross
Abstract: General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (IC...
By Haojian Huang, Zexi Li, Junhao Guo, Yehang Zhang, Wenxuan Peng, Bohan Zhou, Weilin Ruan, Leyi Wu, Chenxu Wang, Jianchong Su, Binghui Xie, Wosong Chen, Yingjie Xu, Tianhao Zhou, Suzeyu Chen, Pukun Zhao, Jiaqi He, Xinyi Li, Runze Li, Peiran Dong, Shaoxiang Dang, Jing Huang, Yingbing Chen, Yifan Chang, Tianyi Zhang, Shiyuan Deng, Haozhi Wang, Yangkai Wei, Wenqian Li, Han Yang, Kaiwen Zhou, Huaping Liu, James Cheng, Rui Shao, Donglin Wang, Yaochu Jin, Jianye Hao, Ying-Cong Chen, Yinchuan Li
arXiv:2609.36390v1 Announce Type: cross
Abstract: Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire functio...
By Gefei Lin, Rui Miao, Xiaoke Zhang
arXiv:2609.37165v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific f...
By Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang
arXiv:2609.37599v1 Announce Type: cross
Abstract: Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators...
By Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson
arXiv:2609.37677v1 Announce Type: cross
Abstract: Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, th...
By Feiyang Wu, Chenxiao Gao, Chen Yang, Ye Zhao, Bo Dai, Anqi Wu
arXiv:2606.08729v2 Announce Type: replace-cross
Abstract: Developing navigation policies requires simulation scenarios that support repeatable training and evaluation. Despite the availability of num...
By Ruihua Han, Shuai Wang, Chengyang Li, Rui Gao, Xinyi Wang, Zhe Liu, Guoliang Li, Yupu Lu, Qi Hao, Jia Pan, Hengshuang Zhao