arXiv AI By Youssef Mahran, Zeyad Gamal, Ayman El-Badawy

Dynamic Entropy Tuning in Reinforcement Learning Low-Level Quadcopter Control: Stochasticity vs Determinism

Read the original on arXiv AI →

arXiv:2512. 18336v2 Announce Type: replace-cross Abstract: This paper explores the impact of dynamic entropy tuning in Reinforcement Learning (RL) algorithms that train a stochastic policy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 15

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

arXiv:2601. 19624v3 Announce Type: replace-cross Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude.

By Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv Machine Learning
Aug 24

Optimistic Online LQR via Intrinsic Rewards

The paper introduces IR‑LQR, an optimistic online linear quadratic regulator that incorporates intrinsic rewards and variance regularization to encourage exploration while maintaining the standard LQR structure. By only adjusting the cost function, IR‑LQR remains computationally simple yet achieves the optimal worst‑case regret rate of √T. The authors validate the method with numerical experiments on aircraft pitch angle control and a UAV example, comparing it to state‑of‑the‑art online LQR algorithms.

By Marcell Bartos, Bruce D. Lee, Lenart Treven, Andreas Krause, Florian D\"orfler, Melanie N. Zeilinger