SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 11119v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models.
arXiv:2602. 02710v2 Announce Type: replace Abstract: Reinforcement learning (RL) is the method of choice for training models in setups where the objective function can only be evaluated by sampling from the model.
Tail-Likelihood Reinforcement Learning (TailRL) is a new approach that optimizes the probability of exceeding randomly chosen reward thresholds instead of just the expected reward. By converting continuous rewards into a family of binary success events, TailRL gives more weight to rare, high-reward rollouts, effectively acting as a mixture of Best‑of‑(k) gradients. The method requires only a simple adjustment to the advantage function, making it compatible with existing reinforcement learning pipelines and improving performance across tasks such as object localization, maze navigation, GUI grounding, and code optimization.
arXiv:2606. 05606v1 Announce Type: new Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed rollout budget for every prompt, despite large differences in the training signal different prompts provide.
arXiv:2606. 01281v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs).
arXiv:2608. 01717v1 Announce Type: new Abstract: Recent reinforcement learning methods for diffusion large language models (dLLMs) commonly rely on on-policy rollouts generated by the target dLLM itself.