Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
The paper formulates and analyzes the linear exponential quadratic Gaussian (LEQG) covariance steering problem in continuous time over a finite horizon. It shows that the optimal controller, still a linear state feedback, cannot be expressed in closed form but is parameterized by a symmetric matrix solving an algebraic equation that captures the risk‑sensitivity parameter. The authors demonstrate that this controller generalizes the risk‑neutral case and prove existence‑uniqueness of solutions near the known risk‑neutral solution for matched noise and input channels, illustrated with a numerical example.
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
arXiv:2507.21543v3 Announce Type: replace-cross Abstract: Mutual information regularization has recently been studied in reinforcement learning (RL) as an extension of entropy regularization, in whic...
arXiv:2606. 28281v1 Announce Type: cross Abstract: PAC-Bayesian bounds provide finite-sample guarantees for data-dependent randomized predictors, but applying them to learning-based control is difficult because the natural objective is a quadratic trajectory cost.
arXiv:2605. 08488v2 Announce Type: replace-cross Abstract: We develop a unified Lyapunov-integral quadratic constraint (IQC) framework for establishing uniform stability of first-order accelerated optimization algorithms in the $\beta$-smooth and $\gamma$-strongly convex regime.
arXiv:2406. 07746v4 Announce Type: replace-cross Abstract: We propose a computationally efficient algorithm that achieves anytime regret of order $\mathcal{O}(\sqrt{t})$, with explicit dependence on the system dimensions and on the solution of the Discrete Algebraic Riccati Equation (DARE).
arXiv:2606. 07600v1 Announce Type: cross Abstract: We formulate data propagation through the Transformer, the machine learning architecture powering large language models, as a nonlinear control system on the space of probability measures.
arXiv:2607. 22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions.
arXiv:2605. 02961v2 Announce Type: replace-cross Abstract: Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-control objective and then learning the associated score or drift field with a neural network.
arXiv:2604. 19569v4 Announce Type: replace-cross Abstract: Q-learning is a fundamental algorithmic primitive in reinforcement learning.
arXiv:2609.07597v1 Announce Type: cross Abstract: Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. This gives a matrix-sign update that preserves the singular...
arXiv:2608. 01151v1 Announce Type: cross Abstract: In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints.
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.