The paper studies how open‑vocabulary segmentation models perform on UAV footage, focusing on temporal consistency of predictions. By linking frame‑wise outputs to persistent 3D voxels via metric fusion, the authors propose a voxel‑level evaluation that measures final agreement, Semantic Belief Drift, Observation Persistence, and uncertainty. Experiments on UAVid‑3D show that high overall agreement can mask instability when observations are sparse, and that persistence‑stratified analysis reveals greater disagreement for recurrent voxels while belief drift reduces with more evidence.
By Saurbh Singh Jamwal
arXiv:2506.21185v3 Announce Type: replace
Abstract: 3D semantic occupancy prediction is crucial for autonomous driving, providing a dense, semantically rich environmental representation. However, exi...
By Yuheng Zhang, Mengfei Duan, Kunyu Peng, Yuhang Wang, Ruiping Liu, Fei Teng, Kai Luo, Zhiyong Li, Kailun Yang
arXiv:2409. 11972v4 Announce Type: replace-cross Abstract: Enabling robots to autonomously discover high-level spatial concepts (e.
By Jose Andres Millan-Romera, Muhammad Shaheer, Miguel Fernandez-Cortizas, Martin R. Oswald, Holger Voos, Jose Luis Sanchez-Lopez
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
By Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He
arXiv:2606. 06721v1 Announce Type: cross Abstract: Robots that operate over extended periods should not merely visit space; they should progressively understand it.
By Junyu Mao, Sara Ayoubi, Vishnu D. Sharma, Ilija Had\v{z}i\'c, Matthew Andrews
arXiv:2608.21136v1 Announce Type: new
Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...
By Jie Xu, Na Zhao