The paper investigates reinforcement learning in Markov decision processes whose dynamics are perturbed by non‑Markovian external events. It identifies conditions that make the problem tractable by limiting consideration to a finite history of events, and proposes a policy iteration algorithm that learns state‑dependent policies conditioned on this history. The authors provide theoretical guarantees for policy improvement, analyze sample complexity for least‑squares evaluation and improvement, and extend their results to discrete‑time Hawkes processes with Gaussian marks, validating their approach with experiments in control environments.
By Ranga Shaarad Ayyagari, Revanth Raj Eega, Ambedkar Dukkipati
arXiv:2506. 08121v2 Announce Type: replace-cross Abstract: We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics.
By Qi Feng, Gu Wang
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2606. 18183v1 Announce Type: cross Abstract: Temporal difference (TD) learning with linear function approximation is a core method for policy evaluation.
By M. Forzo, E. Monzio Compagnoni, A. Russo, A. Pacchiano
arXiv:2608. 01917v1 Announce Type: new Abstract: Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning.
By Ankur Naskar, Vivek T A, Aditya Kumar, Gugan Thoppe, Prashanth L. A
arXiv:2405.08253v4 Announce Type: replace-cross
Abstract: This paper develops a framework for learning in discounted infinite-horizon Markov decision processes (MDPs) with Borel state and action spac...
By Daniel Adelman, Cagla Keceli, Alba V. Olivares-Nadal