Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
arXiv:2606. 04845v1 Announce Type: cross Abstract: Sequential decision-making problems are often modelled as a Markov decision process (MDP).
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
arXiv:2603. 08287v2 Announce Type: replace-cross Abstract: We analyze the Bayesian regret of the Gaussian process posterior sampling reinforcement learning (GP-PSRL) algorithm.
arXiv:2507. 18606v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) provides a principled framework for decision-making in partially observable environments, which can be modeled as Markov decision processes and compactly represented through dynamic decision Bayesian networks.
arXiv:2602. 17086v2 Announce Type: replace-cross Abstract: Dynamic decision-making under model uncertainty is central to many economic environments, yet existing bandit and reinforcement learning algorithms rely on the assumption of correct model specification.
arXiv:2605. 23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.
arXiv:2609.07998v1 Announce Type: new Abstract: We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rathe...
arXiv:2511.20413v2 Announce Type: replace-cross Abstract: \emph{Decision-focused learning} (DFL) trains predictive models to optimize downstream decisions rather than prediction accuracy alone. While...
arXiv:2602. 17375v3 Announce Type: replace Abstract: We formulate episodic Markov decision process (MDP) planning as Bayesian inference over policies.
arXiv:2610.01269v1 Announce Type: cross Abstract: Bayesian Optimisation (BO) is a powerful framework for the optimisation of expensive black-box functions, but typically requires refitting a surrogat...
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
arXiv:2607. 01741v1 Announce Type: cross Abstract: Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environment by maximizing cumulative rewards.
The paper investigates how fast predictive regret guarantees of exact Bayesian online learning can be maintained when using approximate posterior methods. It establishes a general theorem linking the cumulative cost of posterior approximation to the contraction radius of the exact Gibbs posterior and the Wasserstein distance between approximate and exact posteriors. Three concrete online learning scenarios—linear models, infinite‑dimensional exponential families, and Gaussian process regression—illustrate that appropriately accurate approximations (projected Langevin, truncation, and sparse variational posteriors) preserve fast regret bounds while reducing computational demands.