arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang
arXiv:2607. 07737v1 Announce Type: cross Abstract: GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry.
By Natalia Trukhina, Vadim Vashkelis
Metric-Bench introduces a new benchmark for Vision‑Language Models (VLMs) that focuses on metric‑spatial reasoning in indoor scenes by using in‑image reference objects with known dimensions. The accompanying MetricReasoner fine‑tuning recipe employs structured prompts and numerical rewards to implicitly learn 2D‑to‑3D mapping without camera intrinsics. Experiments show that this approach improves spatial metric understanding by 43.1 % over existing models and boosts downstream embodied tasks, while also delivering gains on general VLM benchmarks.
By Yuling Xi, Haokai Zhang, Muzhi Zhu, Hao Zhong, Zongze Du, Hengyu Zhao, Chenchen Jing, Yufei Yin, Bin Qin, Yongjie Yang, Zhenbo Luo, Hao Chen, Chunhua Shen
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
arXiv:2608. 07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments.
By Peng Xu, Chengcheng Wang, Shaohua Wan
The paper introduces a tool‑augmented framework that enhances a small Vision‑Language Model (Qwen3.5‑4B) with geometric tools—3D object detection, metric depth estimation, and deterministic solvers for distance, size, and bearing—to improve metric spatial reasoning. By moving metric computation from the model’s weights into explicit solvers, the approach achieves significant gains on ReVSI‑Bench tasks, notably increasing absolute distance accuracy from 0.46 to 0.74 MRA and relative direction accuracy from 25.9% to 73.4%. The modular design allows swapping in different detectors, enabling a clear separation between perception and reasoning errors, and the model can autonomously sequence the tools to match a scripted pipeline on most tasks.
By Kai Glantz, Clemens Grange