arXiv:2505. 18201v2 Announce Type: replace-cross Abstract: Controlling flapping-wing drones requires controllers that handle time-varying, nonlinear, underactuated dynamics from incomplete, noisy sensor data.
By Romain Poletti, Lorenzo Schena, Lilla Koloszar, Joris Degroote, Miguel Alfonso Mendez
arXiv:2512. 18333v2 Announce Type: replace-cross Abstract: This paper proposes a new Reinforcement Learning (RL) based control architecture for quadrotors.
By Youssef Mahran, Zeyad Gamal, Ayman El-Badawy
The paper proposes a hybrid PID–Deep Reinforcement Learning (DRL) controller for industrial processes, addressing the limitations of traditional PID controllers in complex, non‑linear, multi‑input environments. Using the Industrial Benchmark (IB) to test DRL, the authors develop a multi‑objective reward function and employ a TD3 agent to discover optimal settings for the IB’s ‘Gain’ and ‘Shift’ parameters. These parameters are then fed into a tuned PID controller, yielding a system that combines the optimal performance and efficiency of DRL with the reliability of classical control.
By Zhengyang (Cissy), Gu, Joseph E. Hernandez, John Burtenshaw, Sean Scott, Thomas Cook, Chris Couch
arXiv:2607. 01528v1 Announce Type: new Abstract: Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control.
By Abdullah Al Tasim, Wei Sun
The paper introduces MTD3-BC, a model‑free offline reinforcement learning algorithm that optimizes yaw control for wind farms amid changing wind directions. By learning from a pre‑collected dataset and incorporating an action consistency term, it reduces the need for extensive simulator interactions. Experimental wind‑tunnel tests show that MTD3‑BC improves farm‑level power output by about 10% compared to a greedy baseline and matches a model‑based benchmark, all while cutting training costs dramatically.
By Yuhan Su, Hongyang Dong, Simone Tamaro, Filippo Campagnolo, Carlo L. Bottasso, Xiaowei Zhao
The paper studies recurrent Twin Delayed Deep Deterministic Policy Gradient (TD3) agents in environments with evolving hidden disturbances, focusing on how observation history, action history, history length, and network structure influence performance. Three recurrent architectures are compared under controlled disturbances, revealing that action history is crucial when responses depend on prior actions and that a unified temporal sequence of action-observation pairs outperforms separate branches. The authors introduce H‑TD3, which reuses actor-generated recurrent states to initialize the critic, and demonstrate that these architectures excel in a rover wheel‑slip simulation, with policies trained on temporally structured disturbances transferring better to unseen slip dynamics.
By Saki Omi, Hyo-Sang Shin, Namhoon Cho, Antonios Tsourdos, Miguel A. Olivares-Mendez