arXiv AI

Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief

arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.

arXiv AI
Jul 17

Fully Offline Reinforcement Learning

arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.

By Mattie Fellows, Clarisse Wibault, Uljad Berdica, Johannes Forkel, Maike Osborne, Jakob N. Foerster
arXiv Machine Learning
Jun 25

Efficient Adaptive Data Acquisition via Pretrained Belief Representations

arXiv:2606. 25197v1 Announce Type: new Abstract: Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder.

By Daolang Huang, Zhuoyue Huang, Conor Hassan, Luigi Acerbi, Samuel Kaski, Tom Rainforth
arXiv AI
Sep 4

Subspace Inference Enables Efficient Active Reward Learning from Preferences

The paper introduces PreferenceEKF, a sample‑efficient method for active reward learning from human preferences. By framing preference learning as a sequential Bayesian filtering problem, it tracks reward model uncertainty using an extended Kalman filter in a low‑dimensional subspace, avoiding costly posterior inference over the full neural network. Experiments on D4RL and V‑D4RL benchmarks show improved sample efficiency, runtime, scalability, and calibration, with reward models that support competitive offline reinforcement learning policies.

By Yutai Zhou, Erdem B{\i}y{\i}k
arXiv Machine Learning
Aug 24

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.

By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
arXiv AI
Aug 12

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

arXiv:2605. 23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.

By Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P\'erez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport