Sample complexity of variance-reduced policy gradient: weaker assumptions and lower bounds
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2302.07477v4 Announce Type: replace Abstract: We study the optimal sample complexity of tabular reinforcement learning for infinite-horizon discounted Markov decision processes. The unrestricte...
arXiv:2509. 16586v2 Announce Type: replace Abstract: Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model.
arXiv:2609. 22833v1 Announce Type: new Abstract: We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step.
arXiv:2606. 14095v1 Announce Type: new Abstract: We study the sample complexity of learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits (RBs) under a generative model.
The paper investigates best‑policy identification in finite‑horizon, risk‑sensitive reinforcement learning using the entropic risk measure. It identifies a gap between known lower bounds ≥ η(e^{|eta|H}) and upper bounds ≤ O(e^{2|eta|H}) for sample complexity, attributing the excess factor to loose concentration bounds for exponential utilities. By employing a forward‑model algorithm with KL‑based exploration bonuses and a novel stopping rule, the authors achieve a sample complexity that matches the lower bound, closing the previously open exponential gap.
arXiv:2609.36945v1 Announce Type: new Abstract: We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version...