arXiv AI

Reinforcement Learning under External Influence: Guarantees, Algorithms, and Sample Complexity

The paper investigates reinforcement learning in Markov decision processes whose dynamics are perturbed by non‑Markovian external events. It identifies conditions that make the problem tractable by limiting consideration to a finite history of events, and proposes a policy iteration algorithm that learns state‑dependent policies conditioned on this history. The authors provide theoretical guarantees for policy improvement, analyze sample complexity for least‑squares evaluation and improvement, and extend their results to discrete‑time Hawkes processes with Gaussian marks, validating their approach with experiments in control environments.

arXiv Machine Learning
Aug 20

Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions

The paper introduces a method for controlling multivariate Hawkes-driven stochastic differential equations using machine learning in a non‑Markovian setting. It develops a finite‑dimensional Markovianization technique that approximates Hawkes processes with mixtures of exponential kernels, proving convergence of the approximation to the original process and its value function. A continuous‑time deterministic policy gradient algorithm, called Hawkes‑CT DDPG, is then proposed to solve the control problem model‑free, relying only on observed event times, SDE solutions, and decay filters, and is compared against discrete‑time reinforcement learning approaches for various kernel types.

By Tomasz R. Bielecki, Thibaut Mastrolia, Haoze Yan
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
Hugging Face Trending Papers
Aug 3

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.

arXiv Machine Learning
Jul 7

A Hierarchy of Policy Learning Problems

arXiv:2607. 03385v1 Announce Type: cross Abstract: Policy learning has received substantial attention with the goal of learning policies from observational data for decision-making.

By Hamsa Bastani, Osbert Bastani, Shihan Chen
arXiv Machine Learning
Sep 14

Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics

The paper investigates learning Nash equilibria in partially observable Markov games (POMGs) where agents cannot fully observe the state. By focusing on a subclass with independent state transitions and a Markov potential game structure, the authors propose an independent learning algorithm that allows agents to converge to an approximate Nash equilibrium using only their own observations and actions, without communication. Under a filter stability assumption, finite‑history policies are shown to approximate the POMG sufficiently, enabling a surrogate near‑potential Markov game and yielding quasi‑polynomial sample and computational complexity.

By Philip Jordan, Maryam Kamgarpour
arXiv Machine Learning
Jul 1

End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions

arXiv:2603. 23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Processes (MDPs) satisfying \emph{linear Bellman completeness} -- a fundamental setting where the Bellman backup of any linear value function remains linear.

By Zakaria Mhammedi, Alexander Rakhlin, Nneka Okolo