Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
The paper introduces PreferenceEKF, a sample‑efficient method for active reward learning from human preferences. By framing preference learning as a sequential Bayesian filtering problem, it tracks reward model uncertainty using an extended Kalman filter in a low‑dimensional subspace, avoiding costly posterior inference over the full neural network. Experiments on D4RL and V‑D4RL benchmarks show improved sample efficiency, runtime, scalability, and calibration, with reward models that support competitive offline reinforcement learning policies.
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
arXiv:2607. 19199v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to learn an effective policy from a static dataset, but its performance is fundamentally limited by dataset coverage.
arXiv:2607. 28408v1 Announce Type: new Abstract: This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback.
arXiv:2606. 19328v1 Announce Type: cross Abstract: Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design.
The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.
arXiv:2606. 01123v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) avoids explicit reward engineering by learning from pairwise human preference feedback.
arXiv:2606. 01468v1 Announce Type: cross Abstract: Due to their explicit priors and ability to model uncertainty, Bayesian methods have played a major role in dynamical latent variable modeling of single-cell neural recordings.
arXiv:2505. 22442v3 Announce Type: replace-cross Abstract: Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparameter tuning and lack reliable fully offline estimates of initial online performance.
arXiv:2602. 23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning.
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.