arXiv Machine Learning By Felix St\"orck, Philipp Hartmann, Fabian Hinder, Klaus Neumann, Barbara Hammer

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces Space-sampled Value Decay (SsVD), a forgetting mechanism designed for non-stationary reinforcement learning where the environment can drift at every timestep. SsVD selectively pulls value estimates of randomly chosen state-space elements toward a baseline, discarding outdated information without requiring reset or change-point detection. Integrated into Soft Actor Critic and Deep Q-Networks, SsVD outperforms its base algorithms across six non-stationary environments and can also promote optimism in hard-exploration tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 11

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Deep Reinforcement Learning

arXiv:2606. 11797v1 Announce Type: new Abstract: Studies on rodents such as mice have shown the capabilities to adapt their behavior when dealing with changing parameters (``drift'') of the environment even if no information about change is provided (uncertainty) -- a behavior that can be modeled by forgetting mechanisms.

By Felix St\"orck, Fabian Hinder, Barbara Hammer
arXiv Machine Learning
Sep 17

Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning

The paper investigates how to balance retaining past experience versus learning from new data when robot dynamics change. It introduces two metrics—change magnitude and age‑staleness AUC—to quantify when older transitions are helpful or harmful. Experiments on locomotion tasks and real‑world perturbations show that the optimal replay strategy depends on the size of the dynamics shift and the evolution of the system over time.

By Everest Yang, Skye Thompson, George D. Konidaris