Hugging Face Trending Papers

Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

Read the original on Hugging Face Trending Papers →

Air-Ground Collaborative Vision-and-Language Navigation (AGC-VLN) pairs an unmanned aerial vehicle (UAV) with a global bird’s‑eye view and an unmanned ground vehicle (UGV) with a local first‑person view, creating a shared bird’s‑eye map that enables collaboration. The training‑free baseline decomposes navigation into VLM‑based semantic reasoning and deterministic geometric execution, allowing the UAV to render the UGV’s pose and target markers while the UGV plans a road‑following path using the shared map. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and surpasses the strongest single‑agent baseline by 24.0 points.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 4

Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

Air-Ground Collaborative Vision-and-Language Navigation (AGC‑VLN) pairs a UAV with a global bird’s‑eye view and a UGV with a local first‑person view, creating a shared bird’s‑eye map that displays the UGV’s pose and the target location. The UAV localizes the target in its downward view and flies toward it, while the UGV uses the shared map to plan a road‑following path with a frozen VLM and execute it under closed‑loop control. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and outperforms the strongest single‑agent baseline by 24.0 points.

By Shuning Zhang, Liang Li, Yunheng Wang, Tao Wang, Yihang Kang, Renjing Xu
arXiv AI
Sep 2

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

The paper introduces DroneCATS-Agent, a modular framework that places a multimodal large language model (MLLM) at the core of a drone’s control loop, allowing the model to decide actions solely from prompts. It presents the DroneCATS benchmark, evaluating MLLMs on four tasks—approaching, tracking, searching, and multi‑drone commanding—without fine‑tuning or function‑calling. Results show that while small open models can navigate reliably, they often fail by mismanaging protocol termination, highlighting a gap between perception and action planning in current MLLMs.

By Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu, Daehee Kim, Myungseo Song, Jihyuk Byun, Seunggyu Chang, Taeho Kil, Jiseob Kim, Bado Lee, Geewook Kim
arXiv Machine Learning
Jun 3

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.

By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang
arXiv AI
Jun 19

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

arXiv:2606. 20045v1 Announce Type: cross Abstract: UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly.

By Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun