Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Air-Ground Collaborative Vision-and-Language Navigation (AGC-VLN) pairs an unmanned aerial vehicle (UAV) with a global bird’s‑eye view and an unmanned ground vehicle (UGV) with a local first‑person view, creating a shared bird’s‑eye map that enables collaboration. The training‑free baseline decomposes navigation into VLM‑based semantic reasoning and deterministic geometric execution, allowing the UAV to render the UGV’s pose and target markers while the UGV plans a road‑following path using the shared map. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and surpasses the strongest single‑agent baseline by 24.0 points.
Air-Ground Collaborative Vision-and-Language Navigation (AGC‑VLN) pairs a UAV with a global bird’s‑eye view and a UGV with a local first‑person view, creating a shared bird’s‑eye map that displays the UGV’s pose and the target location. The UAV localizes the target in its downward view and flies toward it, while the UGV uses the shared map to plan a road‑following path with a frozen VLM and execute it under closed‑loop control. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and outperforms the strongest single‑agent baseline by 24.0 points.
arXiv:2606. 14772v1 Announce Type: cross Abstract: Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions.
arXiv:2608. 11738v1 Announce Type: cross Abstract: Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density.
arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.