arXiv AI

Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems

The paper investigates local stationary solutions of finite‑horizon discrete‑time Pontryagin systems near a steady extremal. Under regularity of the stationarity equation, hyperbolicity of the reduced state–costate map, and a scaled transversality condition, the linearized boundary‑value problem admits a uniformly bounded inverse, leading to existence, uniqueness, and uniform Lipschitz estimates independent of the horizon. The study further shows that perturbations of the terminal reward decay exponentially with the horizon, and for linear‑quadratic systems with suitable conditions the Riccati matrix and initial feedback gain converge at a quantified rate, with numerical experiments confirming the theoretical predictions.

arXiv Machine Learning
Jul 28

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.

By Asha Barua, Sajad Khodadadian
arXiv Machine Learning
Aug 17

Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

arXiv:2608. 13651v1 Announce Type: cross Abstract: We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-\Delta, \Delta]$ is unknown in sign and magnitude and $\Delta$ is arbitrarily large.

By Dimitar Ho
arXiv Machine Learning
Jul 21

Scaling Limits of Constant-Stepsize SGD at Flat Minima

arXiv:2607. 16384v1 Announce Type: new Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons.

By Jingyi Zhang, Cheng Mao, Debankur Mukherjee