arXiv Machine Learning

Regret of exploratory policy improvement and $q$-learning

arXiv:2411. 01302v2 Announce Type: replace Abstract: We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J.

arXiv Machine Learning
Aug 4

Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning

arXiv:2608. 02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization.

By Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
arXiv Machine Learning
Jul 7

Adaptive Partitioning and Learning for Stochastic Control of Diffusion Processes

arXiv:2512. 14991v2 Announce Type: replace Abstract: We study reinforcement learning for controlled diffusion processes with unbounded continuous state spaces, bounded continuous actions, and polynomially growing rewards: settings that arise naturally in finance, economics, and operations research.

By Hanqing Jin, Renyuan Xu, Yanzhao Yang