The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.
By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel
arXiv:2606. 01397v1 Announce Type: cross Abstract: A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can degrade another.
By Mehmet Iscan, Batuhan Temiz
This study explores temporal neural networks for estimating the end‑effector position of an aerial continuum manipulator (ACM) affected by aerodynamic disturbances from a UAV. An experimental dataset covering stationary and free‑hovering conditions across various robot configurations and altitudes was used to evaluate strain‑parameterized kinematic models and to benchmark a closed‑form continuous‑time (CfC) neural network against an MLP and a GRU. The CfC network achieved a 22 mm RMSE, outperforming the MLP (36 mm) and GRU (28 mm) by 39.5 % and 20.6 %, respectively, demonstrating the advantage of continuous‑time learning for this task.
By Niloufar Amiri, Houman Masnavi, Farrokh Janabi-Sharifi
arXiv:2607. 03132v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) in industrial control often suffers from lag and overshoot due to purely reactive control based on the current tracking error.
By Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender
The paper introduces MTD3-BC, a model‑free offline reinforcement learning algorithm that optimizes yaw control for wind farms amid changing wind directions. By learning from a pre‑collected dataset and incorporating an action consistency term, it reduces the need for extensive simulator interactions. Experimental wind‑tunnel tests show that MTD3‑BC improves farm‑level power output by about 10% compared to a greedy baseline and matches a model‑based benchmark, all while cutting training costs dramatically.
By Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao
arXiv:2606. 00949v1 Announce Type: cross Abstract: We propose a method combining Multi-Agent Deep Reinforcement Learning (MARL) and eXplainable Deep Learning (XDL) to reduce drag in wall-bounded turbulent flows.
By Federica Tonti, Ricardo Vinuesa