arXiv:2510. 01460v4 Announce Type: replace-cross Abstract: Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online interactions for fine-tuning.
By Lu Li, Tianwei Ni, Yihao Sun, Pierre-Luc Bacon
arXiv:2502. 19544v3 Announce Type: replace Abstract: Leveraging offline data is a promising way to improve the sample efficiency of online reinforcement learning (RL).
By Yi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou, Tianyu Cui, Le Chen, Dieter B\"uchler, Arno Solin, Juho Kannala, Joni Pajarinen
The paper presents an online learning algorithm that significantly boosts data efficiency for reinforcement learning from human feedback (RLHF). It incrementally updates reward and language models as choice data arrives, using a small affirmative nudge, an epistemic neural network for reward uncertainty, and information‑directed exploration. With Gemma LLMs, the method matches offline RLHF trained on 200K labels using fewer than 20K labels, achieving over a 10× improvement in data efficiency, and projects a 1,000× gain when scaled to 1M labels.
By Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy
arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.
By Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster
arXiv:2606. 04305v1 Announce Type: new Abstract: We study online learning with an additional offline dataset in the stochastic linear bandit setting.
By Kushagra Chandak, Toshinori Kitamura, Xiaoqi Tan
arXiv:2510. 19528v2 Announce Type: replace-cross Abstract: We investigate the fundamental problem of leveraging offline data to accelerate online reinforcement learning - a direction with strong potential but limited theoretical grounding.
By Sebastian Reboul, H\'el\`ene Halconruy
arXiv:2602. 17632v3 Announce Type: replace-cross Abstract: Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with value-based RL algorithms typically causes immediate drops in performance.
By Nathan Samuel de Lara, Florian Shkurti
arXiv:2502. 17666v5 Announce Type: replace-cross Abstract: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings.
By Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov
arXiv:2609.35880v1 Announce Type: new
Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertain...
By Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon
arXiv:2607. 11720v1 Announce Type: cross Abstract: Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction.
By Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
arXiv:2506. 06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning.
By Zixuan Dong, Yumi Omori, Keith Ross
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous.