arXiv Machine Learning By Zhaoyu Zhu, Rui Gao, Shuang Li

Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

Read the original on arXiv Machine Learning →

arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 27

Trajectory-Regularized Stochastic Optimal Control via KL Divergence

arXiv:2607. 22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions.

By Mintae Kim, Koushil Sreenath
arXiv Machine Learning
Jun 26

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

arXiv:2606. 26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available.

By Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu