arXiv AI By Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Jun 10

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.

By Heming Zou, Qi Wang, Yun Qu, Yuhang Jiang, Lizhou Cai, Yixiu Mao, Ru Peng, Xin Xu, Weijie Liu, Kai Yang, Saiyong Yang, Xiangyang Ji
arXiv Machine Learning
Aug 21

Maximum Likelihood Reinforcement Learning

arXiv:2602. 02710v2 Announce Type: replace Abstract: Reinforcement learning (RL) is the method of choice for training models in setups where the objective function can only be evaluated by sampling from the model.

By Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Yiding Jiang, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette
arXiv Machine Learning
Sep 4

Tail-Likelihood Reinforcement Learning

Tail-Likelihood Reinforcement Learning (TailRL) is a new approach that optimizes the probability of exceeding randomly chosen reward thresholds instead of just the expected reward. By converting continuous rewards into a family of binary success events, TailRL gives more weight to rare, high-reward rollouts, effectively acting as a mixture of Best‑of‑(k) gradients. The method requires only a simple adjustment to the advantage function, making it compatible with existing reinforcement learning pipelines and improving performance across tasks such as object localization, maze navigation, GUI grounding, and code optimization.

By Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng, Qingyang Wu, Zhongzhu Zhou, Chenfeng Xu, Haiwen Feng, Yuda Song, Aarti Singh, Ruslan Salakhutdinov, J. Andrew Bagnell, Jeff Schneider, Andrea Zanette