arXiv Machine Learning

The Bias of Nonlinear Two-Time-scale Stochastic Approximation under Constant Step-Sizes

arXiv:2609. 20409v1 Announce Type: new Abstract: Two-timescale stochastic approximation (TTSA) is a fundamental tool for analyzing coupled iterative algorithms in reinforcement learning, optimization, and stochastic control.

arXiv Machine Learning
Jul 28

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.

By Asha Barua, Sajad Khodadadian
Hugging Face Trending Papers
Aug 3

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.

arXiv Machine Learning
Aug 27

Non-Asymptotic Bounds for Closed-Loop Identification of Sub-Exponentially Growing Nonlinear Stochastic Systems

The paper studies least squares parameter estimation for discrete‑time, unstable, closed‑loop nonlinear stochastic systems with linearly parametrised uncertainty and additive i.i.d. process noise. By perturbing the control policy with exploratory input and assuming a sub‑exponential input‑to‑state growth property, the authors derive non‑asymptotic bounds on the estimation error whenever the state trajectory remains in an informative region of the state space. When the entire state space is informative, the bounds hold with high probability for all time steps, and the authors illustrate the applicability of their results with examples that extend beyond existing work.

By Seth Siriya, Jingge Zhu, Dragan Ne\v{s}i\'c, Ye Pu