arXiv:2607. 16314v1 Announce Type: cross Abstract: World models, especially based on JEPA architectures, have been shown to learn robust dynamics of various environments.
By Usman M. Khan
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
By Hoang Nguyen, Xiaohao Xu, Xiaonan Huang
arXiv:2606. 18634v1 Announce Type: cross Abstract: To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots.
By Zecheng Yin, Benedict Jun Ma
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
By Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He
arXiv:2608. 15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making.
By Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
arXiv:2608. 14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal.
By Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang
Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover complex dynamic scenarios due to treating all visual tokens uniformly and reasoning with human-selected factors, which lack mechanisms to emphasize task-critical evidence and ignore underlying factors.
arXiv:2607. 08182v1 Announce Type: cross Abstract: Vision-language-action (VLA) models aim to map multimodal inputs to robot actions.
By Qi Lyu, Baicheng Liu, Xudong Wang, Jiahua Dong, Lianqing Liu, Zhi Han
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
By Wancong Zhang, Basile Terver, Artem Zholus, Soham Chitnis, Harsh Sutaria, Mido Assran, Randall Balestriero, Amir Bar, Adrien Bardes, Yann LeCun, Nicolas Ballas
arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.
By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
arXiv:2606. 04226v1 Announce Type: cross Abstract: Simulation environments are useful for both robot policy learning and planning verification and validation.
By Charlie Gauthier, Sacha Morin, Liam Paull
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
By Yuanjie Lu, Beichen Wang, Zhengqi Wu, Yang Li, Xiaomin Lin, Chengzhi Mao, Xuesu Xiao