arXiv:2602. 02250v2 Announce Type: replace-cross Abstract: Kullback-Leibler (KL) divergence regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise regimes.
By Viktor Stein, Adwait Datar, Nihat Ay
arXiv:2506. 08121v2 Announce Type: replace-cross Abstract: We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics.
By Qi Feng, Gu Wang
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2604. 08580v2 Announce Type: replace-cross Abstract: Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints.
By Carles Domingo-Enrich, Jiequn Han
arXiv:2608. 10777v1 Announce Type: new Abstract: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community.
By Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu
arXiv:2602. 05379v2 Announce Type: replace-cross Abstract: Effective reinforcement learning (RL) for complex stochastic systems requires leveraging historical data to improve sample efficiency and accelerate policy optimization.
By Hua Zheng, Wei Xie, M. Ben Feng, Keilung Choy
arXiv:2508. 04225v4 Announce Type: replace-cross Abstract: Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline reinforcement learning.
By Lingwei Zhu, Haseeb Shah, Zheng Chen, Martha White
arXiv:2601. 18840v4 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming.
By Donghwan Lee, Hyukjun Yang
arXiv:2606. 26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available.
By Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2404. 05185v4 Announce Type: replace-cross Abstract: This paper deals with a class of neural SDEs and studies the limiting behavior of the associated sampled optimal control problems as the sample size grows to infinity.
By Huafu Liao, Alp\'ar R. M\'esz\'aros, Chenchen Mou, Chao Zhou
arXiv:2512. 04697v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time reinforcement learning (RL) for optimal switching problems across multiple regimes.
By Yijie Huang, Mengge Li, Xiang Yu, Zhou Zhou