The paper investigates when intrinsic rewards effectively drive exploration in reinforcement learning. It introduces a formal criterion that evaluates policies based on the counterfactual information they acquire, comparing how well their histories can replace experience from alternative policies. Using a simple environment, the authors show that common intrinsic reward objectives—count-based, prediction-error, empowerment, and information-gain—can lead to Pareto-suboptimal exploration under this criterion, and they propose conditions and a new objective that better align with optimal exploration.
By Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)
arXiv:2601. 18840v4 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming.
By Donghwan Lee, Hyukjun Yang
The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.
By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv:2606. 00151v1 Announce Type: cross Abstract: In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal.
By Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
arXiv:2607. 03168v1 Announce Type: cross Abstract: Entropy regularization is widely used in continuous-time reinforcement learning (RL) to reduce sensitivity to environmental perturbations, yet its robustness benefits lack a rigorous theoretical foundation.
By Jialun Cao, Fernando Acero, David \v{S}i\v{s}ka, Yufei Zhang
arXiv:2607. 18077v1 Announce Type: cross Abstract: What gives the Bellman equation its form?
By Fernando E. Rosas, David Hyland, Daniel Polani
arXiv:2607. 17823v1 Announce Type: new Abstract: Reinforcement Learning is a cornerstone technique for modern large reasoning models.
By Riccardo Poiani, Martino Bernasconi, Andrea Celli
arXiv:2601. 05675v2 Announce Type: replace Abstract: Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI.
By Bingyi Liu, Jinbo He, Haiyong Shi, Enshu Wang, Weizhen Han, Jingxiang Hao, Peixi Wang, Zhuangzhuang Zhang
arXiv:2609.39342v1 Announce Type: cross
Abstract: Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful...
By Manolis Mylonas, Rub\'en Moreno Bote
The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
By Sebastian Sanokowski, Kaustubh Patil, Majid Khadiv
The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.
By Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa
arXiv:2607. 08925v1 Announce Type: new Abstract: Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do.
By Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin