arXiv Machine Learning

General Bayesian Policy Learning

arXiv:2602. 23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning.

arXiv AI
Sep 4

Subspace Inference Enables Efficient Active Reward Learning from Preferences

The paper introduces PreferenceEKF, a sample‑efficient method for active reward learning from human preferences. By framing preference learning as a sequential Bayesian filtering problem, it tracks reward model uncertainty using an extended Kalman filter in a low‑dimensional subspace, avoiding costly posterior inference over the full neural network. Experiments on D4RL and V‑D4RL benchmarks show improved sample efficiency, runtime, scalability, and calibration, with reward models that support competitive offline reinforcement learning policies.

By Yutai Zhou, Erdem B{\i}y{\i}k
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv AI
Aug 12

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

arXiv:2605. 23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.

By Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P\'erez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport