arXiv Machine Learning

Semi-Offline Reinforcement Learning for Optimized Text Generation

arXiv:2306. 09712v2 Announce Type: replace Abstract: In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline.

arXiv AI
Aug 24

Efficient Exploration at Scale

The paper presents an online learning algorithm that significantly boosts data efficiency for reinforcement learning from human feedback (RLHF). It incrementally updates reward and language models as choice data arrives, using a small affirmative nudge, an epistemic neural network for reward uncertainty, and information‑directed exploration. With Gemma LLMs, the method matches offline RLHF trained on 200K labels using fewer than 20K labels, achieving over a 10× improvement in data efficiency, and projects a 1,000× gain when scaled to 1M labels.

By Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy
arXiv AI
Jul 17

Fully Offline Reinforcement Learning

arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.

By Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster
arXiv AI
Aug 14

Yes, Q-learning Helps Offline In-Context RL

arXiv:2502. 17666v5 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings.

By Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov