arXiv Machine Learning

Miles v0.1: Production-Level Post-Training

arXiv Machine Learning
Jun 26

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

arXiv:2606. 26997v1 Announce Type: cross Abstract: Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks.

By Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu
arXiv AI
Jun 16

RollArt: Disaggregated Multi-Task Agentic RL Training at Scale

arXiv:2512. 22560v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bound prefill, bandwidth-bound decoding, CPU-heavy environment execution, and bursty reward evaluation.

By Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, Ju Huang, Siran Yang, Yongbin Li, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, Wei Wang
arXiv AI
Sep 7

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

RL-VLA$^3$ is a fully asynchronous distributed reinforcement learning framework designed for Vision‑Language‑Action (VLA) model training. It allows fine‑grained asynchronous interaction between simulation, inference, and training via dynamic batching schedulers and flexible environment sharding, addressing the variable, resource‑intensive latencies of physical simulators. Experiments across multiple simulation backends, VLA architectures, and RL algorithms show throughput gains of up to 85.2% over synchronous baselines while preserving sample efficiency, and the system scales from 8 to 256 GPUs.

By Haoran Sun, Yongjian Guo, Zhong Guan, Shuai Di, Xiaodong Bai, Jing Long, Tianyun Zhao, Mingxi Luo, Hongke Zhao, Likang Wu, Xiaotie Deng, Xu Chu, Xi Xiao, Sheng Wen, Yicheng Gong, Junwu Xiong
arXiv Machine Learning
Sep 23

WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning

WeightBridge is a lightweight library that streamlines weight transfer between trainers and rollout generators in reinforcement learning systems, particularly for large language models. It automatically maps trainer and rollout weight layouts, then performs redundancy‑free, load‑balanced transfers while supporting various synchronization modes. Experiments show that WeightBridge can cut GPU stall time by up to 42× compared to leading open‑source RL frameworks, and it was easily integrated into two different frameworks by a coding agent.

By Xuanlin Jiang, Samuel Hsia, Michael Kuchnik, Zachary DeVito, Minlan Yu, Carole-Jean Wu
arXiv Machine Learning
Aug 12

TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling

arXiv:2608. 10402v1 Announce Type: new Abstract: Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times.

By Yanyu Ren, Xizheng Wang, Xiao Liu, Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, Jie Tang
arXiv Machine Learning
Aug 26

WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

arXiv:2608.24479v1 Announce Type: new Abstract: Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for...

By Zihao Wu, Hongyao Tang, Yi Ma, Huizhong Song, Pengyi Li, Yifu Yuan, Fei Ni, Jinyi Liu, Wei Wei, Jianrong Wang, Yan Zheng, Jianye Hao
arXiv Machine Learning
Aug 10

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents

arXiv:2606. 05597v3 Announce Type: replace Abstract: Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, and trajectories that use more steps and tokens than necessary.

By Hao Bai, Rui Yang, Chenlu Ye, Spencer Whitehead, Aviral Kumar, Tong Zhang
arXiv Machine Learning
Sep 24

EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

EBRL is an asynchronous embodied reinforcement learning training system that overlaps rollout and training stages, pipelines simulation and generation across environment groups, and eliminates synchronization stalls. It employs a fine‑grained resource manager that pools CPU cores and GPU streaming multiprocessors, dynamically adjusting resource quotas and batch sizes based on stage profiles and runtime feedback. Experiments on RLinf with four policies and four simulation benchmarks show EBRL improves rollout throughput by 1.30–3.47× and training convergence by 2.5× over state‑of‑the‑art embodied RL systems.

By Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao