arXiv AI By Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

Read the original on arXiv AI →

LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing computation graphs and activations from rollout for policy updates, or by backpropagating provisional gradients and correcting them later. It introduces two training schedules—LeanGRPO‑Retain and LeanGRPO‑Reweight—that target different model scales and input sizes. Experiments on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan show up to a 1.83× end‑to‑end speedup while preserving the original optimization objective.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 3

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing the same feed-forward backbone for rollout and policy update, thereby avoiding unnecessary gradient tracking. It introduces two training schedules—LeanGRPO‑Retain, which reuses computation graphs and activations, and LeanGRPO‑Reweight, which backpropagates provisional gradients and corrects them later. These methods achieve up to a 1.83× speedup on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan while preserving the original optimization objective.

arXiv AI
Jun 24

Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation

arXiv:2606. 24369v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models (LLMs).

By Sijie Wang, Zhengyu Qing, Zhiqiang Tan, Yiming Yin, Yeqing Zhang, Yaoyuan Wang, Qiang Wang, Xiaowen Chu, Shaohuai Shi