arXiv:2506. 22423v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles (UAVs) depend on onboard sensors for perception, navigation, and control.
By Pritam Dash, Ethan Chan, Nathan P. Lawrence, Karthik Pattabiraman
The paper presents a robust multi‑agent reinforcement learning framework for small unmanned aircraft systems (sUAS) to maintain separation assurance when GPS data is degraded or spoofed. By modeling state observation corruption as a zero‑sum game, the authors derive a closed‑form adversarial perturbation that eliminates iterative inner optimization and can be evaluated in linear time. Integrating this perturbation into a policy‑gradient MARL algorithm yields a counter‑policy that achieves near‑zero collision rates in high‑density simulations even with up to 35% observation corruption, outperforming non‑adversarial baselines.
By Alex Zongo, Filippos Fotiadis, Ufuk Topcu, Peng Wei
ASGARD is a two‑phase teacher‑student framework that protects reinforcement‑learning controllers for UAVs from action‑space attacks. In the teacher phase, an encoder fuses the UAV’s physical state with privileged attack information to generate an action‑attack‑aware latent representation, which trains both the control policy and a monitor that corrects actions before they reach the actuators. The student phase then learns to replicate the encoder and monitor using only the UAV’s physical state history, enabling on‑board resilience. Experiments show that ASGARD remains effective against various attack scenarios, including unseen and stealthy attacks, allowing UAV missions to complete successfully.
By Mohsen Salehi, Karthik Pattabiraman
arXiv:2608. 09467v1 Announce Type: cross Abstract: Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments.
By Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li, Chao Yu, Daxin Tian
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution.
The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.
By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel