UDAV: Uncertainty-Driven Adaptive VLM Waypoint Planner
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty.
arXiv:2608. 12308v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability.
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories.
RiskWorld is a risk‑aware world modeling framework that forecasts shared occupancy and selectively replaces planned trajectories in automated driving. It fuses spatial risk fields, temporal actor context, and visual bird’s‑eye‑view features, using flow‑guided evolution to transport occupancy and signed residuals to correct it. In open‑loop planning on nuScenes, RiskWorld achieves the lowest collision rate over a 3‑second horizon and the second‑best average L2 error, running at 11.5 FPS on a single NVIDIA RTX 4090.
The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.