arXiv Machine Learning By Yukuan Wei, Xudong Li, Lin F. Yang

Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs

Read the original on arXiv Machine Learning →

arXiv:2509. 16586v2 Announce Type: replace Abstract: Recent advances have significantly improved our understanding of the sample complexity of learning in average-reward Markov decision processes (AMDPs) under the generative model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.