arXiv:2302.07477v4 Announce Type: replace
Abstract: We study the optimal sample complexity of tabular reinforcement learning for infinite-horizon discounted Markov decision processes. The unrestricte...
By Shengbo Wang, Jose Blanchet, Peter Glynn
arXiv:2609.36486v1 Announce Type: new
Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M...
By Zijun Chen, Zihan Zhang
arXiv:2509. 16586v2 Announce Type: replace Abstract: Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model.
By Yukuan Wei, Xudong Li, Lin F. Yang
arXiv:2605.15692v2 Announce Type: replace-cross
Abstract: We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that a...
By Zijun Chen, Zihan Zhang
arXiv:2608. 06545v1 Announce Type: new Abstract: Distributionally robust Markov decision processes provide a principled framework for sequential decision making under model uncertainty.
By Yuepeng Yang, Yuxin Chen, Yuejie Chi
arXiv:2606. 09668v1 Announce Type: new Abstract: Contextual queueing bandits provide a framework for learning to schedule heterogeneous jobs under unknown context-dependent service rates.
By Seoungbin Bae, Dabeen Lee