arXiv:2609.23695v1 Announce Type: new
Abstract: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must...
By Mohamed Amine Ferrag, Merouane Debbah, Abderrahmane Lakas, Manu Perumkunnil, Norbert Tihanyi
arXiv:2608. 11738v1 Announce Type: cross Abstract: Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density.
By Haoyu Zhang, Shuoxun Zhang, Peng Ye, Lin Zhang, Jiakang Yuan, Shenghong Yi, Yuening Wang, Tao Chen
The paper proposes Neuro‑Symbolic Agentic AI (NSAAI) as a framework that blends neural grounding, symbolic reasoning, and closed‑loop interaction to enhance decision‑making for networked low‑altitude UAVs. It outlines NSAAI’s strengths in data efficiency, compositional generalization, continual learning, and zero‑shot transfer, and presents a reference architecture covering task management, planning, verification, skill execution, and network interaction. An urban fire‑inspection simulation demonstrates how a UAV can coordinate sensing, cloud access, and verified image‑delivery skills under intermittent connectivity, illustrating NSAAI’s potential for reusable skills, evidence‑grounded decisions, and adaptive mission execution.
By Yuqi Ping, Tianhao Liang, Nanchi Su, Guangyu Lei, Junwei Wu, Qinyu Zhang, Tingting Zhang
The article explores neuro-symbolic agentic AI (NSAAI) as a framework for enhancing the reliability and adaptability of networked low‑altitude UAVs. It outlines NSAAI’s strengths in data efficiency, compositional generalization, continual learning, and zero‑shot transfer, and presents a reference architecture that integrates task management, neuro‑symbolic planning, verification, metacognition, skill execution, and network interaction. A case study of an urban fire‑inspection mission demonstrates how a UAV can coordinate sensing, cloud access, and verified image‑delivery skills under intermittent connectivity, illustrating NSAAI’s potential for reusable skills, evidence‑grounded decision‑making, and adaptive mission execution.
AeroWeaver is a new embodied‑agent harness that integrates large language model (LLM) decision making with the executable skills of individual UAVs, enabling distributed, adaptive swarm execution. It connects semantic mission decisions to governed skills, organizes role‑conditioned local agents for coordination, and refines skill selection online using role‑indexed state‑action‑reward experience. Experiments demonstrate that AeroWeaver maintains valid skill execution without a central joint‑action generator and supports reward‑guided, training‑free adaptive learning from accumulated execution experience.
By Jiabin Lou, Yirong Yang, Haopeng Wang, Xuxin Lv, Xinyu Liu, Diyuan Hou, Xuehong Liu, Rongye Shi, Wenjun Wu
Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinated action. UAV swarms embody...
arXiv:2606. 06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers.
By Xiangyi Zheng, Xiangyu Wang, Qinan Liao, Zimu Tang, Yue Liao, Dongyue Lyu, Guodong Wang, Junjie Liu, Si Liu
arXiv:2608. 15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making.
By Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
arXiv:2606.03963v4 Announce Type: replace-cross
Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but still relies heavily on time consuming manual re...
By Roohan Ahmed Khan, Yasheerah Yaqoot, Amir Atef Habel, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou
arXiv:2607. 17951v1 Announce Type: cross Abstract: Natural-language control offers a promising interface for unmanned aerial vehicles (UAVs), but directly applying self-hosted computer-use agents (SHCUAs) to UAV control introduces a structural mismatch.
By Di Lu, Bo Zhang, Xiyuan Li, Yongzhi Liao, Xuewen Dong, Yulong Shen, Zhiquan Liu, Jianfeng Ma
arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.
By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
Physical Agentic AI proposes an architecture that links semantic planning with physical execution for robot crews. Each robot exposes a typed skill library, while a foundation model planner decomposes tasks into phases and assigns robot‑skill pairs. A Robot Orchestrator validates and authorizes one skill at a time, ensuring actions are grounded in robot capabilities, system state, and workflow constraints before actuation.
By Xinyuan Liu, Eren Sadikoglu, Riana Chatterjee, Ransalu Senanayake