The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.
By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel
arXiv:2606. 01397v1 Announce Type: cross Abstract: A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can degrade another.
By Mehmet Iscan, Batuhan Temiz
This study explores temporal neural networks for estimating the end‑effector position of an aerial continuum manipulator (ACM) affected by aerodynamic disturbances from a UAV. An experimental dataset covering stationary and free‑hovering conditions across various robot configurations and altitudes was used to evaluate strain‑parameterized kinematic models and to benchmark a closed‑form continuous‑time (CfC) neural network against an MLP and a GRU. The CfC network achieved a 22 mm RMSE, outperforming the MLP (36 mm) and GRU (28 mm) by 39.5 % and 20.6 %, respectively, demonstrating the advantage of continuous‑time learning for this task.
By Niloufar Amiri, Houman Masnavi, Farrokh Janabi-Sharifi
arXiv:2607. 03132v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) in industrial control often suffers from lag and overshoot due to purely reactive control based on the current tracking error.
By Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender
The paper introduces MTD3-BC, a model‑free offline reinforcement learning algorithm that optimizes yaw control for wind farms amid changing wind directions. By learning from a pre‑collected dataset and incorporating an action consistency term, it reduces the need for extensive simulator interactions. Experimental wind‑tunnel tests show that MTD3‑BC improves farm‑level power output by about 10% compared to a greedy baseline and matches a model‑based benchmark, all while cutting training costs dramatically.
By Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao
arXiv:2606. 00949v1 Announce Type: cross Abstract: We propose a method combining Multi-Agent Deep Reinforcement Learning (MARL) and eXplainable Deep Learning (XDL) to reduce drag in wall-bounded turbulent flows.
By Federica Tonti, Ricardo Vinuesa
arXiv:2607. 12626v1 Announce Type: cross Abstract: Closed-loop wall control learnt by multi-agent reinforcement learning can lower skin-friction drag in turbulent channels, but these gradient-based policies are trained on small periodic boxes and exhibit reduced performance when carried over to a larger domain.
By Giorgio Maria Cavallazzi, Miguel P\'erez Cuadrado, Alfredo Pinelli
arXiv:2607. 19628v1 Announce Type: new Abstract: In this work we investigate reinforcement learning (RL) as a framework for the robust control of parametrized dynamical systems in presence of measurements and model uncertainties.
By Nicol\`o Botteghi, Gabriele Pascali, Urban Fasel, Andrea Manzoni
BVR Sim is an open‑source, Gymnasium‑style environment for heterogeneous air‑combat reinforcement learning, supporting multiple JSBSim aircraft models (F‑15, F‑16, F/A‑18, F‑22) with configurable weapons, sensors, and opponents. It offers a unified tactical action interface, interchangeable Python and accelerated C++ backends, entity‑oriented observations, compositional rewards, scripted opponents, replay and visualization, and adapters for multi‑agent learning frameworks. At a 0.4‑second decision interval, the C++ backend achieves 104 simulated seconds per wall‑clock second in 1‑vs‑1 and remains practical through 10‑vs‑10 scenarios, and a policy trained on the F‑16 transfers to four unseen aircraft with a 45.5% mean win rate after controller adaptation.
By Haocheng Sun (Beijing University of Posts,Telecommunications), Mulai Tan (Air Force Engineering University)
Closed-loop wall control learnt by multi-agent reinforcement learning can lower skin-friction drag in turbulent channels, but these gradient-based policies are trained on small periodic boxes and exhibit reduced performance when carried over to a larger domain. We recently showed that such policies are also prone to saturated bang-bang actuations that collapse into standing streamwise waves whose scale is set by the computational box rather than by the near-wall cycle, and proposed architectural fixes that avoid these degeneracies.
The paper introduces a correction framework that grounds a CFD-trained deep learning surrogate model for aerospace aerodynamics using wind‑tunnel pressure‑sensor (PSP) data. By training a correction network on spatially registered PSP measurements at two Mach numbers, the authors adjust the surrogate’s predictions without retraining its core parameters, achieving improved agreement with experimental pressure distributions—especially at the wing suction peak and shock location. The grounded surrogate matches measurements within 2.3–2.7% of the Cp range on unseen angles of attack and outperforms simple interpolation between measured states.
By Nitin Nagesh Kulkarni, Dheeraj Vemula, Yin Yu, Peter Lyu, Juan J. Alonso
arXiv:2609. 38977v1 Announce Type: cross Abstract: Neural surrogates have emerged as fast alternatives to the numerical simulation of three-dimensional turbulence.
By Shaoxiang Qin, Yucheng Zhao, Zongyi Li, Liangzhu Leon Wang, Xiongye Xiao