arXiv AI By Jianqiang Xiao, Xiang Deng, Yuexuan Sun, Yanjin Wu, Wenbiao Yan, Liqiang Nie

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

Read the original on arXiv AI →

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots. The method also employs object‑conditioned visual reasoning and egocentric regional guidance, achieving state‑of‑the‑art success rates on the UAV‑ON benchmark.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 8

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots, while also employing object‑conditioned visual reasoning and temporally stable regional guidance. Experiments on the UAV‑ON benchmark show AeroBelief outperforms prior methods in success rate, object success rate, and SPL.

arXiv Computer Vision
Sep 1

Understanding Temporal Semantic Stability in Open-Vocabulary UAV Perception through Metric 3D Fusion

The paper studies how open‑vocabulary segmentation models perform on UAV footage, focusing on temporal consistency of predictions. By linking frame‑wise outputs to persistent 3D voxels via metric fusion, the authors propose a voxel‑level evaluation that measures final agreement, Semantic Belief Drift, Observation Persistence, and uncertainty. Experiments on UAVid‑3D show that high overall agreement can mask instability when observations are sparse, and that persistence‑stratified analysis reveals greater disagreement for recurrent voxels while belief drift reduces with more evidence.

By Saurbh Singh Jamwal
Hugging Face Trending Papers
Jul 13

Parse, Search, and Confirmation: Training-Free Aerial Vision-and-Dialog Navigation with Chain-of-Thought Reasoning and Structured Spatial Memory

In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation. Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.