arXiv Computer Vision By Saurbh Singh Jamwal, Ganesh Ramakrishnan

PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

Read the original on arXiv Computer Vision →

PerSeM is a training‑free framework that builds a persistent semantic memory for long‑horizon UAV mapping by associating frame‑wise segmentation results with world‑space voxels and refining them through spatial refinement, trust‑aware replay, and context‑guided verification. Experiments on Forest and UAVScenes benchmarks show that this persistent 3D memory improves semantic correctness and temporal stability compared to frame‑wise predictions, especially in semantically difficult and temporally unstable regions. The method achieves these gains without retraining or additional neural‑network inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 1

Understanding Temporal Semantic Stability in Open-Vocabulary UAV Perception through Metric 3D Fusion

The paper studies how open‑vocabulary segmentation models perform on UAV footage, focusing on temporal consistency of predictions. By linking frame‑wise outputs to persistent 3D voxels via metric fusion, the authors propose a voxel‑level evaluation that measures final agreement, Semantic Belief Drift, Observation Persistence, and uncertainty. Experiments on UAVid‑3D show that high overall agreement can mask instability when observations are sparse, and that persistence‑stratified analysis reveals greater disagreement for recurrent voxels while belief drift reduces with more evidence.

By Saurbh Singh Jamwal
arXiv AI
Sep 10

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots. The method also employs object‑conditioned visual reasoning and egocentric regional guidance, achieving state‑of‑the‑art success rates on the UAV‑ON benchmark.

By Jianqiang Xiao, Xiang Deng, Yuexuan Sun, Yanjin Wu, Wenbiao Yan, Liqiang Nie
Hugging Face Trending Papers
Sep 8

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots, while also employing object‑conditioned visual reasoning and temporally stable regional guidance. Experiments on the UAV‑ON benchmark show AeroBelief outperforms prior methods in success rate, object success rate, and SPL.

Hugging Face Trending Papers
Aug 19

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem introduces a volatility‑aware memory evolution framework for lifelong scene understanding, combining spatially aligned instance‑level 3D perception with temporal reasoning. It uses a multi‑session SLAM backbone, a reasoning layer that scores evidence and selects memory actions, and a Tri‑Memory structure (Live, Delta, Meta) to preserve current states and event histories. The accompanying LT‑VQA dataset provides multi‑session recordings, persistent identity annotations, and temporal QA pairs, and experiments show LT‑Mem outperforms baselines while using far fewer tokens.