arXiv:2604. 18701v3 Announce Type: replace-cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error across all visited transitions.
By Vin Bhaskara, Haicheng Wang
We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.
The paper introduces Gradient‑Momentum Coupling (GMC), a method that quantifies learning progress by measuring how strongly a sample influences changes in the parameter space, using the normalized absolute product of its gradient and the momentum of previous gradients. GMC filters out noise by accumulating consistent directions of change while canceling random fluctuations, leading to a more uniform prioritization across tasks with varying noise levels and better ranking of learnable tasks by improvement speed. Experiments on MiniGrid MultiRoom tasks show that replacing prediction error with GMC in the Intrinsic Curiosity Module restores exploration capabilities that were lost to unpredictable observations.
By Samuel Blad, Martin L\"angkvist, Amy Loutfi
arXiv:2606. 27766v1 Announce Type: cross Abstract: Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe.
By Shiqiang Gong
The paper introduces DiffusionOPSD, an on‑policy self‑distillation framework that transforms image‑level reinforcement learning rewards into explicit targets for intermediate denoising predictions in diffusion models. By generating trajectories with a frozen behavior policy and constructing bounded positive and negative targets around query states, the method trains a policy to fit these targets before updating the behavior policy via an exponential moving average. Experiments on SD 3.5‑M and Z‑Image‑Turbo show that DiffusionOPSD achieves the best held‑out scores in 19 of 20 reward‑matched settings, outperforms the strongest competitor by up to 44 % and cuts GPU‑hour usage by 40–63 % compared to DiffusionNFT.
By Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, Xian Sun, Yingshuo Wang, Linfeng Li, Shengqiong Wu, Leigang Qu, Feng Li, Wei Liu, Julian McAuley, Tat-Seng Chua
arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.
By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy