arXiv Machine Learning

State-space models through the lens of ensemble control

arXiv Machine Learning
Jun 11

Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity

arXiv:2606. 11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training.

By Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren
arXiv Machine Learning
Aug 26

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

The paper presents a finite‑sample learning‑to‑control framework for geometrically supervised latent models of nonlinear deterministic systems. It introduces an encoder‑only local–global metric hinge that ensures directional resolution and state discrimination, and proves that any approximate empirical minimizer is pointwise co‑Lipschitz and uniformly approximately semiconjugate to the true dynamics under regularity assumptions. The results provide explicit bounds on approximation, sampling, and optimization errors, and demonstrate through controlled experiments that restoring metric resolution improves control performance.

By Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran
arXiv Machine Learning
Jul 3

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control

arXiv:2604. 08580v2 Announce Type: replace-cross Abstract: Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints.

By Carles Domingo-Enrich, Jiequn Han
arXiv Machine Learning
Jul 28

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.

By Asha Barua, Sajad Khodadadian