arXiv Machine Learning

A Bellman Optimality Equation for Plasticity

The paper introduces a Bellman optimality equation for optimizing plasticity in Markov decision processes, extending recent work that reframes the stability‑plasticity dilemma as an empowerment‑plasticity tradeoff. It builds on Abel et al. (2025), which defined plasticity as the generalized directed information from observations to actions and empowerment as the reverse. This work is the first to address plasticity optimization under that new definition, providing a theoretical foundation similar to existing empowerment‑based approaches.

arXiv Machine Learning
1d ago

When Do Intrinsic Rewards Lead to Exploration?

The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.

By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv AI
Jun 2

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.

By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv AI
3d ago

Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning

The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.

By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
arXiv Machine Learning
Sep 3

Reinforcement learning to choose optimizers

The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.

By Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa
arXiv Machine Learning
Jul 13

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.

By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin