Mathematical methods of reinforcement learning
arXiv:2607. 06935v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory.
The paper presents a large deviations framework for efficient data acquisition in infinite-horizon reinforcement learning, introducing the exponential decay rate of policy-selection error probability as a key efficiency metric. It derives a variational characterization leading to a nested optimization problem, then proposes a tractable convex relaxation and a lazy one-step projected subgradient method to construct an adaptive data acquisition policy. The resulting algorithm is shown to be near-robustly optimal under the proposed criterion, with extensions to linear function approximation and supporting numerical experiments.
arXiv:2607. 06935v1 Announce Type: cross Abstract: Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory.
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
arXiv:2608. 01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints.
arXiv:2607. 17823v1 Announce Type: new Abstract: Reinforcement Learning is a cornerstone technique for modern large reasoning models.
arXiv:2608. 03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications.
arXiv:2608. 08644v1 Announce Type: new Abstract: We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution.
We consider the problem of sampling compositional and discrete objects from a given unnormalized posterior distribution. Notably, recent studies have shown that this problem can be efficiently solved by learning a deterministic Markov Decision Process (MDP) that progressively builds each object in proportion to the posterior.
arXiv:2604. 06039v2 Announce Type: replace-cross Abstract: Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL).
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
arXiv:2607. 28408v1 Announce Type: new Abstract: This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback.
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
arXiv:2404. 03578v3 Announce Type: replace Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL).