arXiv Machine Learning By Everest Yang, Skye Thompson, George D. Konidaris

Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper investigates how to balance retaining past experience versus learning from new data when robot dynamics change. It introduces two metrics—change magnitude and age‑staleness AUC—to quantify when older transitions are helpful or harmful. Experiments on locomotion tasks and real‑world perturbations show that the optimal replay strategy depends on the size of the dynamics shift and the evolution of the system over time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 17

Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL

Changepoint-Aware World Models (CAWM) is a DreamerV3 agent that detects abrupt dynamics shifts in a robot’s environment using an online CUSUM test on internal prediction error. Upon detection, CAWM selectively forgets stale replay data while preserving the learned representation, enabling rapid recovery from shifts such as doubled gravity or halved actuator gain. Experiments on simulated locomotion show CAWM recovers faster than passive retraining and outperforms a baseline that respawns a fresh dynamics model, achieving significant return gains in the first 30k post‑shift frames.

By Everest Yang
arXiv Machine Learning
1d ago

Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning

The paper introduces Space-sampled Value Decay (SsVD), a forgetting mechanism designed for non-stationary reinforcement learning where the environment can drift at every timestep. SsVD selectively pulls value estimates of randomly chosen state-space elements toward a baseline, discarding outdated information without requiring reset or change-point detection. Integrated into Soft Actor Critic and Deep Q-Networks, SsVD outperforms its base algorithms across six non-stationary environments and can also promote optimism in hard-exploration tasks.

By Felix St\"orck, Philipp Hartmann, Fabian Hinder, Klaus Neumann, Barbara Hammer
arXiv Machine Learning
Sep 21

Benchmarking World Models for Continual Learning on Compositional Tasks

The paper introduces a compositional continual learning benchmark for world models in robot manipulation, designed to isolate knowledge reuse from learning speed and capacity. Tasks are curated to combine previously seen action and perception components, allowing analysis of how different modalities affect reuse. Experiments show that modular world models better balance reuse and forgetting than conventional methods, yet none fully solve the challenge, highlighting the need for models explicitly built to reuse knowledge without forgetting.

By Haoyu Zhou, Joe Watson, Anson Lei, Ingmar Posner