arXiv Machine Learning By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

Read the original on arXiv Machine Learning →

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.