arXiv AI

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

arXiv:2601. 19624v3 Announce Type: replace-cross Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude.

arXiv Machine Learning
4d ago

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning

The paper introduces Space-sampled Value Decay (SsVD), a forgetting mechanism designed for non-stationary reinforcement learning where the environment can drift at every timestep. SsVD selectively pulls value estimates of randomly chosen state-space elements toward a baseline, discarding outdated information without requiring reset or change-point detection. Integrated into Soft Actor Critic and Deep Q-Networks, SsVD outperforms its base algorithms across six non-stationary environments and can also promote optimism in hard-exploration tasks.

By Felix St\"orck, Philipp Hartmann, Fabian Hinder, Klaus Neumann, Barbara Hammer
arXiv AI
Aug 24

Behavior-Consistent Deep Reinforcement Learning

The paper introduces the concept of behavior-consistent deep reinforcement learning, aiming to produce high-performing policies that remain distributionally similar across different training runs. It shows that maximum-entropy RL can control behavioral divergence by anchoring runs to a common prior, and proves that for Boltzmann policies, a temperature proportional to Q‑function disagreement limits pairwise KL divergence. Building on this, the authors propose Q‑value Expectile Disagreement (QED), a state‑dependent temperature schedule that uses double‑critic disagreement to approximate cross‑run disagreement, and demonstrate that QED reduces across‑run divergence by two orders of magnitude on 18 continuous‑control tasks without sacrificing performance.

By Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv Machine Learning
Sep 23

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy

The paper introduces ELEMENT, a framework that combines episodic and lifelong entropy maximization to drive reward-free exploration in reinforcement learning. It addresses two key limitations of existing entropy-based methods: the vanishing intrinsic reward after a state is visited and the computational cost of estimating entropy over large datasets. ELEMENT achieves this by deriving an average episodic state entropy reward and employing a k‑NN graph‑based estimator for lifelong entropy, leading to superior state coverage and unsupervised pre‑training performance compared to current baselines.

By Hongming Li, Zhao Yang, Xiaoxuan Liang, Shujian Yu, Jose C. Principe
Hugging Face Trending Papers
Aug 20

End-to-end Early Classification of Time Series in Non-Stationary Environments

Early Classification of Time Series (ECTS) requires making accurate decisions as early as possible in inherently online and evolving environments. Yet, most existing methods assume stationarity and rely on separable designs, where classification and triggering are optimized independently, an assumption that fundamentally limits their adaptability under drift.

arXiv Machine Learning
Sep 1

Uncertainty-Driven Replay Memory for Reinforcement Learning

The paper introduces Uncertainty-Driven Replay Memory (UDRM), a new experience replay buffer for reinforcement learning that prioritizes storing transitions with high uncertainty estimates. Unlike traditional buffers that rely on temporal difference error or transition distributions, UDRM updates its contents based on uncertainty derived from the RL model during training. Experiments show that this uncertainty-aware buffer leads to higher rewards during training compared to other uncertainty-aware RL frameworks.

By Sheeraja Rajakrishnan, Alexander G. Ororbia, Travis Desell, Daniel E. Krutz
arXiv AI
Sep 4

Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

Headroom-Drift Replay is a replay control primitive designed for GRPO that separates reuse into two decisions: Headroom ranks stored groups by remaining learning value, while Drift gates them by compatibility with the current policy. The method keeps the fresh on‑policy stream unchanged and adds no auxiliary generation or training machinery. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, it outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32, delivering comparable quality at materially lower wall‑clock time in Agentic Search.

By Hyun Bin Park, Du-Seong Chang