arXiv AI
Sep 10

Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning

The paper presents a policy gradient theorem tailored to Cumulative Prospect Theory (CPT) objectives in finite-horizon reinforcement learning, extending the classic policy gradient framework to include distortion-based risk measures. Leveraging this theorem, the authors develop a first‑order policy gradient algorithm that uses a Monte Carlo estimator based on order statistics, providing statistical guarantees and proving asymptotic convergence to first‑order stationary points of the generally nonconvex CPT objective. The work also offers a non‑asymptotic sample complexity bound for reaching an approximate stationary policy and demonstrates the qualitative effects of CPT through simulations, comparing the new first‑order method to existing zeroth‑order approaches.

By Olivier Lepel, Anas Barakat
arXiv Machine Learning
Sep 3

Cantelli Constrained Policy Optimization

The paper introduces Canary, a risk‑averse reinforcement learning method that optimizes Value‑at‑Risk (VaR) constraints. By applying Cantelli’s inequality, Canary derives a tractable, conservative, and smooth bound on the VaR constraint using only the first two moments of the cost return, yielding a stable constraint estimator even with tight violation thresholds. Extending the trust‑region framework of Constrained Policy Optimization (CPO), the authors provide worst‑case bounds for policy improvement and constraint violation, and empirically demonstrate that Canary reliably satisfies the VaR constraint in every tested environment.

By Rohan Tangri, Jan-Peter Calliess