arXiv:2607. 22201v1 Announce Type: cross Abstract: We introduce trajectory-regularized stochastic optimal control (TRSOC), which augments standard stochastic optimal control (SOC) with a Kullback--Leibler (KL) divergence between controlled and reference trajectory distributions.
By Mintae Kim, Koushil Sreenath
arXiv:2601. 18840v4 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming.
By Donghwan Lee, Hyukjun Yang
arXiv:2608. 12111v1 Announce Type: cross Abstract: A novel advective Fisher-Rao metric is introduced for optimization tasks on paths of probability measures governed by the continuity equation.
By Benjamin Gess, Johannes M\"uller
arXiv:2608. 12828v1 Announce Type: cross Abstract: Distribution steering seeks feedback laws that drive the state law of a dynamical system between prescribed initial and terminal distributions.
By Kaito Ito, Anqi Dong
arXiv:2608. 02487v1 Announce Type: cross Abstract: Recently, rectified flow has emerged as a fundamental framework for large-scale image generation, powering state-of-the-art systems such as FLUX.
By Leda Wang, Zhehao Xu, Qiang Liu, Harrison H. Zhou
arXiv:2605. 26078v3 Announce Type: replace Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action distributions.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2604.24226v3 Announce Type: replace-cross
Abstract: In this paper we study continuum-marginal optimal transport. Given a time-continuous family of probability marginals, the problem is to recov...
By Yumiharu Nakano
arXiv:2607. 25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.
By Debjyoti Paul
We study optimal-policy geometry in structured Markov decision processes. While approximate dynamic programming and reinforcement learning typically approximate high-dimensional value functions, we show that optimal policies induce simpler decision tessellations.
arXiv:2606. 21585v4 Announce Type: replace Abstract: A finite agent, a machine's digital twin or any bounded reasoner, infers a fixed and noisy world through finite sensors, so its coherent output is a belief: a probability density over states (the Bayes posterior).
By Laurent Caraffa
arXiv:2608. 07433v1 Announce Type: cross Abstract: Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space.
By Zhaoyu Zhu, Rui Gao, Shuang Li
arXiv:2606. 09047v1 Announce Type: cross Abstract: A classical universal stabilization formula offers the practitioner no design freedom: it is a single, parameter-free object.
By Miroslav Krstic, Luke Bhan