arXiv AI By Sen Wang, Yiming Sun, Jiaxuan He, Pengfei Zhu

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

Read the original on arXiv AI →

The paper introduces RACO, a reliability‑aware adaptive coarse‑to‑fine navigation framework for inspection‑oriented UAV vision‑language navigation. It treats the coarse goal as a runtime hypothesis, using object‑level anchors to correct localization before and at the transition to the fine stage, and applies scale‑adaptive terminal refinement for near‑miss cases. RACO is evaluated on the new LG‑UVI inspection setting and outperforms the HETT baseline by 9.53 and 7.98 percentage points on validation‑unseen and test‑unseen, respectively, while improving inspection‑region arrival and reducing false verification risk.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.

By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
arXiv AI
Jun 19

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

arXiv:2606. 20045v1 Announce Type: cross Abstract: UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly.

By Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun
Hugging Face Trending Papers
Jun 18

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its field of view.

Hugging Face Trending Papers
Jul 21

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing approaches to vision-language navigation (VLN) for Unmanned Aerial Vehicle (UAV) typically relies on a single inference pass, which can falter in complex environments by producing suboptimal or unsafe trajectories.