arXiv Machine Learning

A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.

Hugging Face Trending Papers
Aug 3

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.