arXiv:2608.29296v1 Announce Type: cross
Abstract: Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this stati...
By Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen
arXiv:2606. 03070v1 Announce Type: cross Abstract: Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optimization, but stale responses introduce distribution drift.
By Zehua Liu, Yuxuan Yao, Xiaojin Fu, Tao Zhong, Mingxuan Yuan
Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint.
arXiv:2607. 18722v1 Announce Type: new Abstract: Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing.
By Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv:2606. 11087v1 Announce Type: cross Abstract: Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simulated and real robot control.
By Zhiyuan Zhou, Andy Peng, Charles Xu, Qiyang Li, Tobias Springenberg, Kevin Frans, Sergey Levine
arXiv:2605. 05481v2 Announce Type: replace Abstract: We revisit a classic "chicken-and-egg" problem in reinforcement learning: to safely improve a policy, the value function must be accurate on the state-visitation distribution of the updated policy.
By Dillon Sandhu, Ronald Parr
arXiv:2610.01133v1 Announce Type: cross
Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yiel...
By Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu
arXiv:2609.06421v1 Announce Type: cross
Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies repor...
By Daniel Palenicek, Mikael Henaff, Scott Fujimoto, Koustuv Sinha
arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.
By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay. Through controlled experimen...
The paper introduces the concept of behavior-consistent deep reinforcement learning, aiming to produce high-performing policies that remain distributionally similar across different training runs. It shows that maximum-entropy RL can control behavioral divergence by anchoring runs to a common prior, and proves that for Boltzmann policies, a temperature proportional to Q‑function disagreement limits pairwise KL divergence. Building on this, the authors propose Q‑value Expectile Disagreement (QED), a state‑dependent temperature schedule that uses double‑critic disagreement to approximate cross‑run disagreement, and demonstrate that QED reduces across‑run divergence by two orders of magnitude on 18 continuous‑control tasks without sacrificing performance.
By Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker, Benjamin Eysenbach, Eric Eaton
arXiv:2607. 20822v1 Announce Type: new Abstract: Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback.
By Sreejeet Maity, Aritra Mitra