arXiv Machine Learning

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

RollVerify is a lightweight reinforcement learning framework that addresses the trade‑off between efficiency and accuracy in long‑tail rollout settings. It introduces an off‑policy shift metric (OPS) to quantify deviation in partially generated trajectories and uses sequence‑level and token‑level verification to truncate invalid suffixes before training. Experiments on mathematical reasoning tasks show that RollVerify matches on‑policy accuracy while cutting training costs, with preliminary evidence from code‑generation tasks.

arXiv Machine Learning
Jul 21

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

arXiv:2604. 26256v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-tailed trajectories indispensable for model performance block the entire training pipeline.

By Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hongyu Zang, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi Gu, Yerui Sun, Yuchen Xie, Xunliang Cai
Hugging Face Trending Papers
4d ago

ThunderSyncRL: Lossless Acceleration of Agentic Reinforcement Learning

ThunderSyncRL is a training framework that eliminates idle time in agentic reinforcement learning by starting gradient computation immediately once all necessary inputs are available, thereby avoiding policy staleness. It applies to group relative policy optimization (GRPO) by computing trajectory score gradients as soon as rewards arrive, and to on‑policy distillation (OPD) by updating gradients for completed agentic turns while tool calls execute. Experiments on SWE‑bench Verified and Terminal Bench 4.0 show that ThunderSyncRL matches synchronous training performance up to 1.9× faster and outperforms asynchronous training by up to 2.47 percentage points at a fixed budget.