arXiv:2607. 02037v1 Announce Type: cross Abstract: Autonomous surface vehicles vary widely in hydrodynamic and actuation characteristics, yet most controllers are designed for single-platform deployment.
By Ruiheng Jiang, Thomas Bi, Raffaello D'Andrea, Aswin Ramachandran
arXiv:2505. 08222v3 Announce Type: replace-cross Abstract: Autonomous vehicles (AVs) offer a cost-effective solution for scientific missions such as underwater tracking.
By Matteo Gallici, Ivan Masmitja, Mario Mart\'in
arXiv:2606. 08513v1 Announce Type: cross Abstract: Autonomous Underwater Vehicles (AUVs) traditionally rely on complex, heavily engineered pipelines for perception, path planning, and motion control.
By Elisei Shafer, Oren Gal
arXiv:2603. 29426v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning (MARL) provides a promising solution for cooperative target tracking in networks of autonomous underwater vehicles (AUVs).
By Jiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Zhenyu Wang, Chen An
arXiv:2608.20948v1 Announce Type: cross
Abstract: Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this...
By Zhitao Liu, Guangtong Xu, Zihan Wang, Jialiang Hou, Chao Xu, Fei Gao
arXiv:2608. 14135v1 Announce Type: cross Abstract: Autonomous pursuit-evasion is a fundamental challenge for Unmanned Aerial Vehicles (UAVs), requiring rapid decision-making under tightly coupled dynamics and continuously changing opponent behaviors.
By Wenhao Tang, Tianyang Chen, Zhejun Cui, Boyuan An, Jiayu Chen, Ruize Zhang, Huidong Liu, Tianyue Wu, Qingmin Liao, Fei Gao, Yu Wang, Chao Yu
The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.
By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel
arXiv:2604. 12645v2 Announce Type: replace-cross Abstract: Although autonomous underwater vehicles promise the capability of marine ecosystem monitoring, their deployment is fundamentally limited by the difficulty of controlling vehicles under highly uncertain and non-stationary underwater dynamics.
By Melvin Laux, Yi-Ling Liu, Rina Alo, S\"oren T\"opper, Mariela De Lucas Alvarez, Frank Kirchner, Rebecca Adam
arXiv:2607. 16313v1 Announce Type: new Abstract: Traditional controllers are designed for specific systems and do not transfer across different system orders and dynamics.
By Klinsmann Agyei, Pouria Sarhadi
The paper presents a method where self‑evolving scientific agents design explicit, neural‑network‑free white‑box controllers for fluid dynamics tasks. By iteratively interpreting simulation data, the agents accumulate control knowledge and refine controller code, ultimately achieving robust control of an underactuated two‑joint swimmer in unsteady flows. The resulting controllers generalize across varying target positions, wake geometries, cylinder counts, and inflow speeds, and 2D control priors successfully transfer to accelerate 3D adaptation.
By Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
The paper introduces Sampling-Guided Policy Search (SGPS), a method that combines sampling-based model‑predictive control with first‑order policy gradients to accelerate visual policy learning for locomotion and manipulation tasks. SGPS starts with behavior cloning from sampled actions and then alternates between sampling‑based refinement and short‑horizon policy updates under varied initial states and dynamics. The approach is demonstrated on simulated Unitree Go2 and G1 robots, learning tasks such as obstacle traversal and bimanual carrying, and the distilled policies transfer zero‑shot to a real Go2 robot using onboard depth perception.
By Yilang Liu, Haoxiang You, Qian Wang, Daniel Rakita, Ian Abraham
Safe quadrotor navigation in cluttered and dynamic environments depends not only on instantaneous geometric perception, but more critically on anticipating collision risks induced by relative motion. Conventional modular pipelines frequently suffer from perception latency, while end-to-end learning methods relying on implicit scalar rewards often struggle to extract reliable spatio-temporal features without physics-grounded supervision.