Hugging Face Trending Papers

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty.

Hugging Face Trending Papers
Sep 8

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots, while also employing object‑conditioned visual reasoning and temporally stable regional guidance. Experiments on the UAV‑ON benchmark show AeroBelief outperforms prior methods in success rate, object success rate, and SPL.

arXiv AI
Sep 10

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots. The method also employs object‑conditioned visual reasoning and egocentric regional guidance, achieving state‑of‑the‑art success rates on the UAV‑ON benchmark.

By Jianqiang Xiao, Xiang Deng, Yuexuan Sun, Yanjin Wu, Wenbiao Yan, Liqiang Nie
arXiv AI
Jun 15

Schr\"odinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation

arXiv:2512. 21201v3 Announce Type: replace-cross Abstract: Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-built maps, a key capability for general-purpose service robots.

By Yu He, Da Huang, Zhenyang Liu, Zixiao Gu, Qiang Sun, Guangnan Ye, Yanwei Fu, Yu-Gang Jiang
arXiv Computer Vision
Sep 18

Feeling Terrain Before Crossing: World Models for Off-Road Navigation

The paper introduces Feel‑WM, an off‑road navigation world model that incorporates proprioceptive data to predict both visual scenes and the robot’s physical sensations such as slip, tilt, and shake. By learning a future proprioceptive state and failure risk from the robot’s own experience, the model can evaluate planned trajectories using a separable score that balances goal similarity with predicted failure risk. Experiments on real and simulated off‑road data show that Feel‑WM outperforms visual‑only models in both open‑loop planning and closed‑loop navigation for wheeled and legged robots, and it successfully guides a Husky robot around rough terrain on mountain trails where an end‑to‑end policy fails.

By E-In Son, Dong-Wook Kim, Ji-Hoon Hwang, Kangsun Lee, Jisung Bae, Jung-Taak Kim, Seung-Woo Seo
Hugging Face Trending Papers
Jul 21

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories.

arXiv AI
Aug 28

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.

By Zengmao Wang, Wei Gao, Shuhan Shen