arXiv AI

Dynamics Models for Offline Hyperparameter Selection in Real-World RL

arXiv:2608. 11349v1 Announce Type: cross Abstract: A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unavailable and online experimentation is costly.

arXiv AI
Jul 17

Fully Offline Reinforcement Learning

arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.

By Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster
arXiv AI
Jul 16

Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners

arXiv:2607. 13274v1 Announce Type: cross Abstract: Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited.

By Haseeb Shah, Lingwei Zhu, Adam White, Martha White
arXiv AI
Aug 14

Yes, Q-learning Helps Offline In-Context RL

arXiv:2502. 17666v5 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings.

By Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov
arXiv Machine Learning
Sep 18

Improving Generalization and Robustness in Offline Reinforcement Learning via Boundary-Aware Data Augmentation

The paper introduces BADA, a Boundary-Aware Data Augmentation technique for offline reinforcement learning. By interpolating neighboring states to create synthetic data that respects the original distribution, BADA improves in-distribution generalization and robustness. Experiments on limited offline datasets show that BADA achieves state-of-the-art performance across diverse benchmarks.

By Gong Gao, Weidong Zhao, Xianhui Liu
arXiv AI
Sep 4

Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs

The paper introduces DR‑Gym, an open‑source, Gymnasium‑compatible environment that simulates electric utility demand‑response programs at the market level. It uses a regime‑switching wholesale price model calibrated to real extreme events and physics‑based building demand profiles, providing a rich observational space and a configurable multi‑objective reward function for reinforcement learning. Baseline strategies and data snapshots demonstrate the simulator’s realism and learnability.

By Jose E. Aguilar Escamilla, Lingdong Zhou, Xiangqi Zhu, Huazheng Wang