arXiv AI By Jan-Niklas Klein, Sona Ghahremani, Christian Medeiros Adriano, Holger Giese

CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation

Read the original on arXiv AI →

arXiv:2606. 16935v1 Announce Type: cross Abstract: Rovers rely on perception to maintain spatial maps that encode both objects and sensor quality (e.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 18

PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

PerSeM is a training‑free framework that builds a persistent semantic memory for long‑horizon UAV mapping by associating frame‑wise segmentation results with world‑space voxels and refining them through spatial refinement, trust‑aware replay, and context‑guided verification. Experiments on Forest and UAVScenes benchmarks show that this persistent 3D memory improves semantic correctness and temporal stability compared to frame‑wise predictions, especially in semantically difficult and temporally unstable regions. The method achieves these gains without retraining or additional neural‑network inference.

By Saurbh Singh Jamwal, Ganesh Ramakrishnan
arXiv AI
Jun 16

MapDream: Task-Driven Map Learning for Vision-Language Navigation

arXiv:2602. 00222v3 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception.

By Guoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang, Maiyue Chen, Kaihui Wang, Bo Zhang, Zhizhong Su, Deying Li, Zhaoxin Fan
arXiv AI
6d ago

NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation

NavGen introduces a text-to-video data generation pipeline that creates about 400K vision‑language navigation episodes for both indoor and outdoor scenes, using high‑fidelity visual generative models. The approach includes a style‑diversification method to scale up rare, hard‑to‑collect data. Models trained on NavGen data outperform those trained on existing UAV navigation datasets and achieve a 75% success rate in real‑world flying experiments.

By Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu, Yijin Wang, Zhiyang Liu, Fei Gao, Yuze Wu, Xin Zhou
arXiv AI
Aug 10

LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation

arXiv:2608. 07079v1 Announce Type: cross Abstract: Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigation and cross-floor navigation are still commonly addressed separately.

By Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen