arXiv:2607. 22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse.
By Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen
The paper introduces SORL, a framework that stabilizes off‑policy reinforcement learning for long‑horizon large language model agents. It identifies two key instability sources—token‑level policy granularity mismatches and high‑variance off‑policy updates—and proposes turn‑level importance sampling and clipping‑triggered normalization to align optimization with multi‑turn interactions. Two instantiations, SO‑PPO and SO‑GRPO, are evaluated on open‑domain, multi‑hop, and medical QA benchmarks, as well as on asynchronous RL for mathematical reasoning, showing improved robustness without the need for early stopping or heuristic tuning.
By Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong
arXiv:2606. 29526v1 Announce Type: new Abstract: Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse.
By Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, Ju Huang, Wenbo Su, Jinyi Liu, Yan Zheng, Jianye Hao, Bo Zheng
arXiv:2605.12070v3 Announce Type: replace-cross
Abstract: Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy o...
By Zhong Guan, Yongjian Guo, Haoran Sun, Wen Huang, Shuai Di, Likang Wu, Xiong Jun Wu, Hongke Zhao
BRACE introduces an anchored Bellman‑residual correction to address stale critic bias in asynchronous reinforcement learning for language models. By limiting the correction horizon to a prefix of policy tokens and adding a constant‑weight Monte‑Carlo tail, it separates policy correction from reward propagation. The method improves mean@1 on BrowseComp‑Plus by 2.4% and runs 2.46× faster per step than synchronous training while staying stable 50 updates off‑policy.
By Guanqun Zhao, Zijun Xie, Binbin Zheng, Jiafeng Lu, Enlei Gong, Zeyu Chen
arXiv:2608.20909v1 Announce Type: new
Abstract: Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled l...
By Xuyao Lin, Yixiang Shan, Jinru Duan, Tao Yang, Xinyu Zhao, Runyu Lei, Yiming Zhao, Jiaxin Fan, Zongbao Feng, Peng Jia
arXiv:2602. 04879v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) serving as the de facto standard algorithm.
By Penghui Qi, Xiangxin Zhou, Zichen Liu, Tianyu Pang, Chao Du, Min Lin, Wee Sun Lee
arXiv:2607. 18830v1 Announce Type: cross Abstract: Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning global policy parameters that can be rapidly adapted to new tasks.
By Garvit Singla, Uma Maheswari Natarajan, Raghuram Bharadwaj Diddigi
arXiv:2607. 07508v1 Announce Type: cross Abstract: Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs).
By Zhenyu Hou, Yujiang Li, Jie Tang, Yuxiao Dong
arXiv:2609.36830v1 Announce Type: new
Abstract: Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation wi...
By Chenliang Li, Neiwen Ling, Zijun Wei, Alfredo Garcia
arXiv:2609.37119v1 Announce Type: cross
Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
By Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry
Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on...