arXiv:2607. 01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution.
By Sung June Kim, Sangpil Kim, Honglak Lee
On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semantic mismatch between the executed visual stream and the original language instruction.
arXiv:2608. 06128v1 Announce Type: new Abstract: Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning.
By Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
arXiv:2606. 12550v1 Announce Type: cross Abstract: Open-world mapless navigation from sparse language instructions requires resolving underspecified goals and inferring which environmental cues are relevant for reaching the goal.
By Arthur Zhang, Carl Qi, Donne Su, Xiangyun Meng, Amy Zhang, Joydeep Biswas
arXiv:2608.31005v1 Announce Type: new
Abstract: Existing long-video agents acquire evidence through one uniform behavior, ignoring whether the required evidence is concentrated, requires broad occurr...
By Can Zhang, Baofeng Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Shuohuan Wang, Dianhai Yu, Ruirui Li
arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.
By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
arXiv:2606. 29892v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning.
By Siyao Chen, Jiakang Yuan, Jiaxin Wang, Tao Chen
arXiv:2609.15606v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on short video understanding yet remain limited on long videos due to the...
By Weixin Xu, Zhenyu Yang, Bing Wang, Shengsheng Qian, Changsheng Xu
arXiv:2609.39915v1 Announce Type: new
Abstract: Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents...
By Haoxiang Shi, Zaijing Li, Muhe Ding, Xiang Deng, Yaowei Wang, Liqiang Nie
AVERT-VLN introduces a closed‑loop framework for vision‑and‑language navigation that incorporates an abstention‑aware Monitor to detect instruction‑execution inconsistencies. The Monitor is trained on a new LOSTNAV dataset of 20K counterfactual risk trajectories and fine‑tuned on 40K normal trajectories to recognize semantic deviations. During deployment, the Monitor can suspend autonomous navigation and request human guidance, while offline preference learning uses deviation‑associated failures to improve the policy, achieving 76.2% and 66.3% success on R2R‑CE and RxR‑CE unseen splits.
By Minrui Liu, Jingke Wang, Yuehao Huang, Hao Su, Jiajun Lv, Yukai Ma, Yong Liu
arXiv:2609.24576v1 Announce Type: cross
Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...
By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv:2604. 17473v3 Announce Type: replace-cross Abstract: Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions.
By Kangyi Wu, Pengna Li, Kailin Lyu, Xi Lin, Lin Zhao, Qingrong He, Jinjun Wang, Jianyi Liu