arXiv Machine Learning By Mingyi Li, Taira Tsuchiya, Kenji Yamanishi

Data- and Variance-dependent Regret Bounds for Online Tabular MDPs

Read the original on arXiv Machine Learning →

arXiv:2602. 01903v2 Announce Type: replace Abstract: This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent regret bounds in the adversarial regime and variance-dependent regret bounds in the stochastic regime.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.