The paper introduces Space-sampled Value Decay (SsVD), a forgetting mechanism designed for non-stationary reinforcement learning where the environment can drift at every timestep. SsVD selectively pulls value estimates of randomly chosen state-space elements toward a baseline, discarding outdated information without requiring reset or change-point detection. Integrated into Soft Actor Critic and Deep Q-Networks, SsVD outperforms its base algorithms across six non-stationary environments and can also promote optimism in hard-exploration tasks.
By Felix St\"orck, Philipp Hartmann, Fabian Hinder, Klaus Neumann, Barbara Hammer
arXiv:2608. 09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time.
By Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
arXiv:2608. 20044v1 Announce Type: new Abstract: Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments.
By Aur\'elien Renault, Alexis Bondu, Antoine Cornu\'ejols, Vincent Lemaire
arXiv:2512. 18336v2 Announce Type: replace-cross Abstract: This paper explores the impact of dynamic entropy tuning in Reinforcement Learning (RL) algorithms that train a stochastic policy.
By Youssef Mahran, Zeyad Gamal, Ayman El-Badawy
The paper introduces the concept of behavior-consistent deep reinforcement learning, aiming to produce high-performing policies that remain distributionally similar across different training runs. It shows that maximum-entropy RL can control behavioral divergence by anchoring runs to a common prior, and proves that for Boltzmann policies, a temperature proportional to Q‑function disagreement limits pairwise KL divergence. Building on this, the authors propose Q‑value Expectile Disagreement (QED), a state‑dependent temperature schedule that uses double‑critic disagreement to approximate cross‑run disagreement, and demonstrate that QED reduces across‑run divergence by two orders of magnitude on 18 continuous‑control tasks without sacrificing performance.
By Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton
arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.
By Soichiro Nishimori, Paavo Parmas