arXiv Machine Learning By Tomasz R. Bielecki, Thibaut Mastrolia, Haoze Yan

Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

Read the original on arXiv Machine Learning →

The paper introduces a method for controlling multivariate Hawkes-driven stochastic differential equations using machine learning in a non‑Markovian setting. It develops a finite‑dimensional Markovianization technique that approximates Hawkes processes with mixtures of exponential kernels, proving convergence of the approximation to the original process and its value function. A continuous‑time deterministic policy gradient algorithm, called Hawkes‑CT DDPG, is then proposed to solve the control problem model‑free, relying only on observed event times, SDE solutions, and decay filters, and is compared against discrete‑time reinforcement learning approaches for various kernel types.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 21

Reinforcement Learning under External Influence: Guarantees, Algorithms, and Sample Complexity

The paper investigates reinforcement learning in Markov decision processes whose dynamics are perturbed by non‑Markovian external events. It identifies conditions that make the problem tractable by limiting consideration to a finite history of events, and proposes a policy iteration algorithm that learns state‑dependent policies conditioned on this history. The authors provide theoretical guarantees for policy improvement, analyze sample complexity for least‑squares evaluation and improvement, and extend their results to discrete‑time Hawkes processes with Gaussian marks, validating their approach with experiments in control environments.

By Ranga Shaarad Ayyagari, Revanth Raj Eega, Ambedkar Dukkipati
Hugging Face Trending Papers
Aug 3

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.