arXiv:2606. 11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training.
By Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren
arXiv:2608. 19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored.
By Zhiqiang Tan
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2505. 11602v3 Announce Type: replace Abstract: Selective State-Space Models (SSMs) such as Mamba have become central to long-sequence modeling.
By Nikola Zubi\'c, Davide Scaramuzza
The paper presents a finite‑sample learning‑to‑control framework for geometrically supervised latent models of nonlinear deterministic systems. It introduces an encoder‑only local–global metric hinge that ensures directional resolution and state discrimination, and proves that any approximate empirical minimizer is pointwise co‑Lipschitz and uniformly approximately semiconjugate to the true dynamics under regularity assumptions. The results provide explicit bounds on approximation, sampling, and optimization errors, and demonstrate through controlled experiments that restoring metric resolution improves control performance.
By Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran
arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
arXiv:2604. 06039v2 Announce Type: replace-cross Abstract: Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL).
By Zhichao Jia, Guanghui Lan
arXiv:2607. 19628v1 Announce Type: new Abstract: In this work we investigate reinforcement learning (RL) as a framework for the robust control of parametrized dynamical systems in presence of measurements and model uncertainties.
By Nicol\`o Botteghi, Gabriele Pascali, Urban Fasel, Andrea Manzoni
arXiv:2604. 08580v2 Announce Type: replace-cross Abstract: Reward fine-tuning of diffusion and flow models and sampling from tilted or Boltzmann distributions can both be formulated as stochastic optimal control (SOC) problems, where learning an optimal generative dynamics corresponds to optimizing a control under SDE constraints.
By Carles Domingo-Enrich, Jiequn Han
arXiv:2607. 12360v1 Announce Type: new Abstract: The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others.
By Subham Singh, Ashutosh Mishra, Subha Raut
arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.
By Asha Barua, Sajad Khodadadian
arXiv:2607. 20769v1 Announce Type: new Abstract: Learning-enabled decision systems often use offline data or computation to reduce online compute cost.
By Shijie Pan, Agustin Castellano, Zeyu Shen, Enrique Mallada