arXiv:2609.36502v1 Announce Type: cross
Abstract: Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platf...
By Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Ashish Gurung, Ishan Miglani, Shivang Gupta, Zachary Levonian, Conrad Borchers, Kenneth R. Koedinger
arXiv:2609.36525v1 Announce Type: cross
Abstract: Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Altho...
By Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan
arXiv:2609.36588v1 Announce Type: cross
Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLA...
By Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li
arXiv:2609.36562v1 Announce Type: cross
Abstract: While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal im...
By Ruochen Zhang, Yao Huang, Yitong Sun, Jiahe Xie, Jin Yan, Jifan Ma, Yuanfang Guo, Xingxing Wei
arXiv:2609.36645v1 Announce Type: cross
Abstract: Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encou...
By Hanseul Kim, Jewon Yeom, Youngjoon Jeong, Minsoo Jo, Taesup Kim
arXiv:2609.36808v1 Announce Type: cross
Abstract: Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the...
By Long Li, Qichao Zhao, Yue Yang, Fan Xu, Zhe Wang, Alan Wee-Chung Liew, Chao Qu, Heng Tao Shen, Shirui Pan
arXiv:2609.36915v1 Announce Type: cross
Abstract: Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities f...
By Rui Huang, Yanlin Mu, Lidong Li, Yucong Wang, Zichen Yan, Lin Zhao
arXiv:2609.36995v1 Announce Type: cross
Abstract: Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-relate...
By Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin, Yushi Huang, Xin Zhang, Yi Zhang, Yu Liu, Jun Zhang
arXiv:2609.37170v1 Announce Type: cross
Abstract: Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation tr...
By Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao, Jing Lyu, Dong Liu
arXiv:2609.37243v1 Announce Type: cross
Abstract: Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typica...
By Dae Ung Jo, Jongin Lim, YoungJoon Yoo, Daeho Um
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2609.37519v1 Announce Type: cross
Abstract: Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched...
By Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang, Tolga Izdas, Zahra Shahrooei, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh
arXiv:2609.37591v1 Announce Type: cross
Abstract: Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time...
By Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang, Ziju Han, Yahong Han
arXiv:2609.37617v1 Announce Type: cross
Abstract: Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target....
By Yunzhe Li, Kyoungjun Park, Hongzi Zhu, Lili Qiu
arXiv:2609.37709v1 Announce Type: cross
Abstract: Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not...
By Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta
arXiv:2609.37871v1 Announce Type: cross
Abstract: Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDr...
By Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou, Lisheng Wu, Xuewei Li, Zequn Qin, Xi Li
arXiv:2602.11678v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual understanding, yet they suffer from a critical limitation: struct...
By Chengwei Ma, Zhen Tian, Zhou Zhou, Zhixian Xu, Xiaowei Zhu, Xia Hua, Si Shi, F. Richard Yu
arXiv:2609.34653v2 Announce Type: replace
Abstract: On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: t...
By Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu, Xuanzhe Liu
arXiv:2506.11261v2 Announce Type: replace-cross
Abstract: Vision-language-action (VLA) models have shown promising progress in robotic manipulation. However, directly mapping visual observations and...
By Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid
arXiv:2508.13744v3 Announce Type: replace-cross
Abstract: Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when...
By Yeji Park, Minyoung Lee, Sanghyuk Chun, Junsuk Choe