arXiv Machine Learning By Mingjie Hu, Jian-Qiang Hu, Enlu Zhou

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

Read the original on arXiv Machine Learning →

The paper presents a large deviations framework for efficient data acquisition in infinite-horizon reinforcement learning, introducing the exponential decay rate of policy-selection error probability as a key efficiency metric. It derives a variational characterization leading to a nested optimization problem, then proposes a tractable convex relaxation and a lazy one-step projected subgradient method to construct an adaptive data acquisition policy. The resulting algorithm is shown to be near-robustly optimal under the proposed criterion, with extensions to linear function approximation and supporting numerical experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 24

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.

By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes