arXiv:2605. 08488v2 Announce Type: replace-cross Abstract: We develop a unified Lyapunov-integral quadratic constraint (IQC) framework for establishing uniform stability of first-order accelerated optimization algorithms in the $\beta$-smooth and $\gamma$-strongly convex regime.
By Don Li, Dacian Daescu
arXiv:2602. 04132v4 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has achieved remarkable success in solving complex sequential decision-making problems.
By Dhruv S. Kushwaha, Zoleikha A. Biron
arXiv:2608. 10777v1 Announce Type: new Abstract: Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community.
By Bangyan Liao, Chenglei Yu, Yuchen Yang, Chuanrui Wang, Zhisheng Song, Peidong Liu, Tailin Wu
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2608. 01917v1 Announce Type: new Abstract: Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning.
By Ankur Naskar, Vivek T A, Aditya Kumar, Gugan Thoppe, Prashanth L. A
arXiv:2606. 15978v1 Announce Type: new Abstract: Tsitsiklis proved convergence of Monte Carlo optimistic policy iteration under a uniform update structure and identified nonuniform update frequencies as a delicate obstruction.
By Yuanlong Chen
arXiv:2606. 09047v1 Announce Type: cross Abstract: A classical universal stabilization formula offers the practitioner no design freedom: it is a single, parameter-free object.
By Miroslav Krstic, Luke Bhan
arXiv:2603. 08287v2 Announce Type: replace-cross Abstract: We analyze the Bayesian regret of the Gaussian process posterior sampling reinforcement learning (GP-PSRL) algorithm.
By Hamish Flynn, Joe Watson, Ingmar Posner, Jan Peters
arXiv:2601. 23229v2 Announce Type: replace Abstract: Markov decision processes (MDPs) are a fundamental model in sequential decision making.
By Ali Asadi, Krishnendu Chatterjee, Ehsan Goharshady, Mehrdad Karrabi, Alipasha Montaseri, Carlo Pagano
arXiv:2606. 28281v1 Announce Type: cross Abstract: PAC-Bayesian bounds provide finite-sample guarantees for data-dependent randomized predictors, but applying them to learning-based control is difficult because the natural objective is a quadratic trajectory cost.
By Domagoj Herceg
arXiv:2608. 13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large.
By Dimitar Ho
arXiv:2607. 02050v1 Announce Type: new Abstract: Motivated by the challenge of stabilizing a general unknown linear dynamical system (LDS) from observations, we study the natural prerequisite of online prediction.
By Yuval Ran-Milo, Angelos Assos, Elad Hazan