arXiv Machine Learning By Masahiro Kato

General Bayesian Policy Learning

Read the original on arXiv Machine Learning →

arXiv:2602. 23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

Subspace Inference Enables Efficient Active Reward Learning from Preferences

The paper introduces PreferenceEKF, a sample‑efficient method for active reward learning from human preferences. By framing preference learning as a sequential Bayesian filtering problem, it tracks reward model uncertainty using an extended Kalman filter in a low‑dimensional subspace, avoiding costly posterior inference over the full neural network. Experiments on D4RL and V‑D4RL benchmarks show improved sample efficiency, runtime, scalability, and calibration, with reward models that support competitive offline reinforcement learning policies.

By Yutai Zhou, Erdem B{\i}y{\i}k
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu