RLLBC-Lib is an educational code library designed to lower the entry barrier for students learning reinforcement learning (RL) in the context of learning-based control. It offers a comprehensive collection of tabular RL methods to reinforce theoretical foundations, followed by a deep RL library that mirrors the same design principles to highlight parallels between simple and state‑of‑the‑art approaches. The library also includes implementations that illustrate core RL principles, contrast RL with other learning‑based control methods, and serve as a foundation for programming assignments with automated grading.
By Bernd Frauenknecht, Emma Cramer, Artur Eisele, Paul Kruse, Lukas Kesper, Jonas Hertrampf, Ramil Sabirov, Jyotirmaya Patra, Johannes Berger, Paul Brunzema, Friedrich Solowjow, Sebastian Trimpe
Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing. However, RL builds on rather complex interactions between different objects that play out over s...
arXiv:2201. 05000v3 Announce Type: replace-cross Abstract: Reinforcement Learning and, recently, Deep Reinforcement Learning are popular methods for solving sequential decision-making problems modeled as Markov Decision Processes.
By Reza Refaei Afshar, Joaquin Vanschoren, Uzay Kaymak, Rui Zhang, Yaoxin Wu, Wen Song, Yingqian Zhang
The paper introduces a set of software packages in R, Python, Julia, and C++ that solve the Sorted L-One Penalized Estimation (SLOPE) problem efficiently. The packages employ a hybrid coordinate descent algorithm capable of fitting generalized linear models with various loss functions such as Gaussian, binomial, Poisson, and multinomial logistic regression. They support dense, sparse, and out‑of‑memory data structures, can compute the full SLOPE path, perform cross‑validation (including relaxed SLOPE), and are shown to outperform existing SLOPE implementations in speed on both real and simulated data.
By Johan Larsson, Malgorzata Bogdan, Krystyna Grzesiak, Mathurin Massias, Jonas Wallin
arXiv:2602. 07832v3 Announce Type: replace-cross Abstract: Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking.
By Xian Wu, Kaijie Zhu, Ying Zhang, Lun Wang, Wenbo Guo