Tail-Likelihood Reinforcement Learning (TailRL) is a new approach that optimizes the probability of exceeding randomly chosen reward thresholds instead of just the expected reward. By converting continuous rewards into a family of binary success events, TailRL gives more weight to rare, high-reward rollouts, effectively acting as a mixture of Best‑of‑(k) gradients. The method requires only a simple adjustment to the advantage function, making it compatible with existing reinforcement learning pipelines and improving performance across tasks such as object localization, maze navigation, GUI grounding, and code optimization.
By Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar, Guanning Zeng, Qingyang Wu, Zhongzhu Zhou, Chenfeng Xu, Haiwen Feng, Yuda Song, Aarti Singh, Ruslan Salakhutdinov, J. Andrew Bagnell, Jeff Schneider, Andrea Zanette
arXiv:2605. 20256v2 Announce Type: replace Abstract: Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models.
By Xikai Zhang, Yongzhi Li, Likang Xiao, Yingze Zhang, Yanhua Cheng, Quan Chen, Peng Jiang, Wenjun Wu, Liu Liu
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
arXiv:2607. 17823v1 Announce Type: new Abstract: Reinforcement Learning is a cornerstone technique for modern large reasoning models.
By Riccardo Poiani, Martino Bernasconi, Andrea Celli
arXiv:2609.36552v1 Announce Type: cross
Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite r...
By Zihao Chen, Fanxiang Xiong, Hongran Ren, Xuefeng Bai, Zhongxiang Dai, Kehai Chen, Zhiguo Zhang, Zhiyong Wang, Yu Cheng
The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.
By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen