arXiv Machine Learning By Giorgio Maria Cavallazzi, Miguel P\'erez-Cuadrado, Alfredo Pinelli

Reward hacking in physical reinforcement learning revealed by turbulent drag reduction

Read the original on arXiv Machine Learning →

arXiv:2606. 06227v2 Announce Type: replace-cross Abstract: A reinforcement-learning agent maximises its reward, which can diverge from the outcome its designer intended.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 15

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

arXiv:2607. 12626v1 Announce Type: cross Abstract: Closed-loop wall control learnt by multi-agent reinforcement learning can lower skin-friction drag in turbulent channels, but these gradient-based policies are trained on small periodic boxes and exhibit reduced performance when carried over to a larger domain.

By Giorgio Maria Cavallazzi, Miguel P\'erez Cuadrado, Alfredo Pinelli
Hugging Face Trending Papers
Jul 14

Gradient-free learning of a closed-loop wall controller for turbulent drag reduction

Closed-loop wall control learnt by multi-agent reinforcement learning can lower skin-friction drag in turbulent channels, but these gradient-based policies are trained on small periodic boxes and exhibit reduced performance when carried over to a larger domain. We recently showed that such policies are also prone to saturated bang-bang actuations that collapse into standing streamwise waves whose scale is set by the computational box rather than by the near-wall cycle, and proposed architectural fixes that avoid these degeneracies.

arXiv AI
Jun 16

QPILOTS: Efficient Test-Time Q-Steering for Flow Policies

arXiv:2606. 14801v1 Announce Type: cross Abstract: Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult.

By Yifan Ruan, Chenyang Cao, Andreas Burger, Ali Pesaranghader, Kaveh Kamali, Jaehong Kim, Nandita Vijaykumar, Alan Aspuru-Guzik, Igor Gilitschenski, Nicholas Rhinehart