arXiv:2604.09107v2 Announce Type: replace-cross
Abstract: Modern LLM reinforcement learning (RL) workloads require a high-performance weight transfer system to scale training across heterogeneous com...
By Chenhao Ye, Huaizheng Zhang, Mingcong Han, Baoquan Zhong, Xiang Li, Qixiang Chen, Xinyi Zhang, Weidong Zhang, Kaihua Jiang, Wang Zhang, He Sun, Wencong Xiao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
arXiv:2512. 22560v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation.
By Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, Ju Huang, Siran Yang, Yongbin Li, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, Wei Wang
arXiv:2609.08368v1 Announce Type: new
Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...
By RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng
arXiv:2606. 26997v1 Announce Type: cross Abstract: Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks.
By Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu
arXiv:2609.24456v1 Announce Type: cross
Abstract: Distributed reinforcement learning (RL) scales training by parallelizing actors and learners around an Experience Buffer. As RL workloads grow, howev...
By Sitong Zhang, Tuo Shi, Mario Di Francesco, Zeke Wang, Bo Zhao
arXiv:2608.22167v1 Announce Type: new
Abstract: Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL framework...
By Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew, Silvio Savarese, Junnan Li
EBRL is an asynchronous embodied reinforcement learning training system that overlaps rollout and training stages, pipelines simulation and generation across environment groups, and eliminates synchronization stalls. It employs a fine‑grained resource manager that pools CPU cores and GPU streaming multiprocessors, dynamically adjusting resource quotas and batch sizes based on stage profiles and runtime feedback. Experiments on RLinf with four policies and four simulation benchmarks show EBRL improves rollout throughput by 1.30–3.47× and training convergence by 2.5× over state‑of‑the‑art embodied RL systems.
By Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao
arXiv:2608. 06025v1 Announce Type: new Abstract: In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations.
By Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu
RL-VLA$^3$ is a fully asynchronous distributed reinforcement learning framework designed for Vision‑Language‑Action (VLA) model training. It allows fine‑grained asynchronous interaction between simulation, inference, and training via dynamic batching schedulers and flexible environment sharding, addressing the variable, resource‑intensive latencies of physical simulators. Experiments across multiple simulation backends, VLA architectures, and RL algorithms show throughput gains of up to 85.2% over synchronous baselines while preserving sample efficiency, and the system scales from 8 to 256 GPUs.
By Haoran Sun, Yongjian Guo, Zhong Guan, Shuai Di, Xiaodong Bai, Jing Long, Tianyun Zhao, Mingxi Luo, Hongke Zhao, Likang Wu, Xiaotie Deng, Xu Chu, Xi Xiao, Sheng Wen, Yicheng Gong, Junwu Xiong
arXiv:2608. 11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms.
By Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian
arXiv:2606. 03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to build, synthetic training queries are often detached from the server's actual state (so the generated tool calls fail to execute), and recall-based RL rewards incentivize verbose tool-calling patterns.
By Ibrahim Abdelaziz, Asim Munawar, Kinjal Basu, Maxwell Crouse, Chulaka Gunasekara, Suneet Katrekar, Pavan Kapanipathi
arXiv:2606. 03077v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a standard post-training paradigm for large language models (LLMs), extending beyond preference alignment to complex reasoning and multi-turn agentic behaviors.
By Kaiwen Chen, Xin Tan, Jingzong Li, Hong Xu