Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective
Read the original on arXiv Machine Learning →The paper presents a large deviations framework for efficient data acquisition in infinite-horizon reinforcement learning, introducing the exponential decay rate of policy-selection error probability as a key efficiency metric. It derives a variational characterization leading to a nested optimization problem, then proposes a tractable convex relaxation and a lazy one-step projected subgradient method to construct an adaptive data acquisition policy. The resulting algorithm is shown to be near-robustly optimal under the proposed criterion, with extensions to linear function approximation and supporting numerical experiments.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.