Hugging Face Trending Papers

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing the same feed-forward backbone for rollout and policy update, thereby avoiding unnecessary gradient tracking. It introduces two training schedules—LeanGRPO‑Retain, which reuses computation graphs and activations, and LeanGRPO‑Reweight, which backpropagates provisional gradients and corrects them later. These methods achieve up to a 1.83× speedup on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan while preserving the original optimization objective.

arXiv AI
Sep 4

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing computation graphs and activations from rollout for policy updates, or by backpropagating provisional gradients and correcting them later. It introduces two training schedules—LeanGRPO‑Retain and LeanGRPO‑Reweight—that target different model scales and input sizes. Experiments on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan show up to a 1.83× end‑to‑end speedup while preserving the original optimization objective.

By Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi
arXiv AI
Jun 24

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

arXiv:2606. 24369v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models (LLMs).

By Sijie Wang, Zhengyu Qing, Zhiqiang Tan, Yiming Yin, Yeqing Zhang, Yaoyuan Wang, Qiang Wang, Xiaowen Chu, Shaohuai Shi
arXiv AI
Aug 3

RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving

arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.

By Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy
arXiv AI
Sep 18

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design

The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.

By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen
arXiv Machine Learning
Jul 21

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

arXiv:2604. 26256v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-tailed trajectories indispensable for model performance block the entire training pipeline.

By Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hongyu Zang, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi Gu, Yerui Sun, Yuchen Xie, Xunliang Cai