arXiv Machine Learning By Abdullah Al Tasim, Wei Sun

Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence

Read the original on arXiv Machine Learning →

arXiv:2607. 01528v1 Announce Type: new Abstract: Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 14

Curriculum-Based Adversarial Heterogeneous Agent Reinforcement Learning for Autonomous Quad-Copter Landing in Maritime Settings

The paper presents a curriculum‑based adversarial heterogeneous agent reinforcement learning (HARL‑AC) approach for autonomous quad‑copter landing on a ship deck in maritime settings. Using Heterogeneous‑Agent Proximal Policy Optimization (HAPPO) in NVIDIA Isaac Lab, the authors train a cooperative control policy that outperforms domain‑randomized baselines, achieving up to 97.5% success on in‑distribution sea states and higher median success and lower crash rates on out‑of‑distribution sea states. The adversarially trained policy also exhibits more cautious behavior, slightly increasing timeouts but improving safety in severe, unseen conditions.

By Allan Minh-Tam Nguyen, Sree Showrya Kotala, Stefan Banioi-Crijman, Kurt Driessens, Rico M\"ockel
arXiv AI
Sep 25

Temporal Learning for End-Effector Position Estimation under Aerodynamic Disturbances in Aerial Continuum Manipulation

This study explores temporal neural networks for estimating the end‑effector position of an aerial continuum manipulator (ACM) affected by aerodynamic disturbances from a UAV. An experimental dataset covering stationary and free‑hovering conditions across various robot configurations and altitudes was used to evaluate strain‑parameterized kinematic models and to benchmark a closed‑form continuous‑time (CfC) neural network against an MLP and a GRU. The CfC network achieved a 22 mm RMSE, outperforming the MLP (36 mm) and GRU (28 mm) by 39.5 % and 20.6 %, respectively, demonstrating the advantage of continuous‑time learning for this task.

By Niloufar Amiri, Houman Masnavi, Farrokh Janabi-Sharifi
arXiv Machine Learning
Sep 14

Offline Reinforcement Learning for Wind Farm Control: A Wind Tunnel Study under Dynamic Wind Directions

The paper introduces MTD3-BC, a model‑free offline reinforcement learning algorithm that optimizes yaw control for wind farms amid changing wind directions. By learning from a pre‑collected dataset and incorporating an action consistency term, it reduces the need for extensive simulator interactions. Experimental wind‑tunnel tests show that MTD3‑BC improves farm‑level power output by about 10% compared to a greedy baseline and matches a model‑based benchmark, all while cutting training costs dramatically.

By Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao