arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2602. 12643v2 Announce Type: replace-cross Abstract: We present Unified Latent Dynamics (ULD), a novel reinforcement learning algorithm that unifies the efficiency of model-free methods with the representational strengths of model-based approaches, without incurring planning overhead.
By Jashaswimalya Acharjee, Balaraman Ravindran
The paper introduces the concept of behavior-consistent deep reinforcement learning, aiming to produce high-performing policies that remain distributionally similar across different training runs. It shows that maximum-entropy RL can control behavioral divergence by anchoring runs to a common prior, and proves that for Boltzmann policies, a temperature proportional to Q‑function disagreement limits pairwise KL divergence. Building on this, the authors propose Q‑value Expectile Disagreement (QED), a state‑dependent temperature schedule that uses double‑critic disagreement to approximate cross‑run disagreement, and demonstrate that QED reduces across‑run divergence by two orders of magnitude on 18 continuous‑control tasks without sacrificing performance.
By Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton
The paper introduces BADA, a Boundary-Aware Data Augmentation technique for offline reinforcement learning. By interpolating neighboring states to create synthetic data that respects the original distribution, BADA improves in-distribution generalization and robustness. Experiments on limited offline datasets show that BADA achieves state-of-the-art performance across diverse benchmarks.
By Gong Gao, Weidong Zhao, Xianhui Liu
arXiv:2606. 20411v1 Announce Type: new Abstract: Direct Advantage Estimation (DAE) has been shown to improve the sample efficiency of deep reinforcement learning algorithms.
By Hsiao-Ru Pan, Bernhard Sch\"olkopf