PerSeM is a training‑free framework that builds a persistent semantic memory for long‑horizon UAV mapping by associating frame‑wise segmentation results with world‑space voxels and refining them through spatial refinement, trust‑aware replay, and context‑guided verification. Experiments on Forest and UAVScenes benchmarks show that this persistent 3D memory improves semantic correctness and temporal stability compared to frame‑wise predictions, especially in semantically difficult and temporally unstable regions. The method achieves these gains without retraining or additional neural‑network inference.
By Saurbh Singh Jamwal, Ganesh Ramakrishnan
arXiv:2607. 16366v1 Announce Type: cross Abstract: Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazards such as steep slopes and rocky terrain.
By Raul Castilla-Arquillo, Carlos Perez-del-Pulgar, Levin Gerdes, Alfonso Garcia-Cerezo, Miguel A. Olivares-Mendez
arXiv:2602. 00222v3 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception.
By Guoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang, Maiyue Chen, Kaihui Wang, Bo Zhang, Zhizhong Su, Deying Li, Zhaoxin Fan
NavGen introduces a text-to-video data generation pipeline that creates about 400K vision‑language navigation episodes for both indoor and outdoor scenes, using high‑fidelity visual generative models. The approach includes a style‑diversification method to scale up rare, hard‑to‑collect data. Models trained on NavGen data outperform those trained on existing UAV navigation datasets and achieve a 75% success rate in real‑world flying experiments.
By Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu, Yijin Wang, Zhiyang Liu, Fei Gao, Yuze Wu, Xin Zhou
arXiv:2608. 07079v1 Announce Type: cross Abstract: Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately.
By Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen
arXiv:2607. 07737v1 Announce Type: cross Abstract: GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry.
By Natalia Trukhina, Vadim Vashkelis
arXiv:2606. 31919v1 Announce Type: cross Abstract: Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of explicit depth information introduces severe physical uncertainty and semantic-physical misalignment.
By Wenyuan Xie, Shaokai Wu, Yijin Zhou, Yanbiao Ji, Guodong Zhang, Bayram Bayramli, Qiuchang Li, Xunchu Zhou, Yue Ding, Hongtao Lu
Mem2Ego introduces a vision‑language model for embodied navigation that combines global memory with egocentric visual inputs. By adaptively retrieving task‑relevant cues from a global memory module and aligning them with local perception, the framework improves spatial reasoning and decision‑making over long horizons. The method outperforms prior state‑of‑the‑art approaches on the HSSD and HM3D benchmarks and shows strong performance on a real robot.
By Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang, Matin Aghaei, Yixin Xiao, Yaochen Hu, Mohammad Ali Alomrani, David Gamaliel Arcos Bravo, Hongjian Gu, Zhiyuan Li, Yangzheng Wu, Zhanpeng Zhang, Raika Karimi, Atia Hamidizadeh, Guowei Huang, Haoping Xu, Tongtong Cao, Weichao Qiu, Xingyue Quan, Jianye Hao, Yuzheng Zhuang, Yingxue Zhang
arXiv:2605. 06317v4 Announce Type: replace-cross Abstract: Existing Vision-Language Navigation (VLN) methods typically adopt an egocentric, step-by-step paradigm, which struggles with error accumulation and limits efficiency.
By Dijia Zhan, Jinyi Li, Chenxi Zheng, Shaoyu Huang, Yong Li, Jie Tang, Xuemiao Xu
arXiv:2602.02220v3 Announce Type: replace
Abstract: Language-conditioned goal navigation (LGN) requires embodied agents to locate user-specified targets without step-by-step guidance. However, existi...
By Bo Miao, Weijia Liu, Jun Luo, Lachlan Shinnick, Jian Liu, Thomas Hamilton-Smith, Yuhe Yang, Zijie Wu, Vanja Videnovic, Feras Dayoub, Anton van den Hengel
GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry. Cross-view image retrieval can provide such fixes, but raw appearance is sensitive to season, illumination, viewpoint, map age, and sensor modality.
arXiv:2609.15098v1 Announce Type: new
Abstract: Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and execut...
By Jianhe Zhao, Yanhua Qiu, Zhiyu Zhang, Zibo Zhao, Jinhua Xie