Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
Read the original on arXiv AI →The paper presents a policy gradient theorem tailored to Cumulative Prospect Theory (CPT) objectives in finite-horizon reinforcement learning, extending the classic policy gradient framework to include distortion-based risk measures. Leveraging this theorem, the authors develop a first‑order policy gradient algorithm that uses a Monte Carlo estimator based on order statistics, providing statistical guarantees and proving asymptotic convergence to first‑order stationary points of the generally nonconvex CPT objective. The work also offers a non‑asymptotic sample complexity bound for reaching an approximate stationary policy and demonstrates the qualitative effects of CPT through simulations, comparing the new first‑order method to existing zeroth‑order approaches.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.