On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events.
arXiv:2601. 22993v4 Announce Type: replace Abstract: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
The paper presents a policy gradient theorem tailored to Cumulative Prospect Theory (CPT) objectives in finite-horizon reinforcement learning, extending the classic policy gradient framework to include distortion-based risk measures. Leveraging this theorem, the authors develop a first‑order policy gradient algorithm that uses a Monte Carlo estimator based on order statistics, providing statistical guarantees and proving asymptotic convergence to first‑order stationary points of the generally nonconvex CPT objective. The work also offers a non‑asymptotic sample complexity bound for reaching an approximate stationary policy and demonstrates the qualitative effects of CPT through simulations, comparing the new first‑order method to existing zeroth‑order approaches.
The paper introduces Canary, a risk‑averse reinforcement learning method that optimizes Value‑at‑Risk (VaR) constraints. By applying Cantelli’s inequality, Canary derives a tractable, conservative, and smooth bound on the VaR constraint using only the first two moments of the cost return, yielding a stable constraint estimator even with tight violation thresholds. Extending the trust‑region framework of Constrained Policy Optimization (CPO), the authors provide worst‑case bounds for policy improvement and constraint violation, and empirically demonstrate that Canary reliably satisfies the VaR constraint in every tested environment.
arXiv:2609.24103v1 Announce Type: new Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
arXiv:2608. 03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications.