AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2603. 12893v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learning from reward signals to explicitly improve desirable aspects such as image quality and prompt alignment.
arXiv:2509.25050v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where both pre-training and RL post-train...
The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.
arXiv:2608. 14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards.
LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing computation graphs and activations from rollout for policy updates, or by backpropagating provisional gradients and correcting them later. It introduces two training schedules—LeanGRPO‑Retain and LeanGRPO‑Reweight—that target different model scales and input sizes. Experiments on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan show up to a 1.83× end‑to‑end speedup while preserving the original optimization objective.
LeanGRPO eliminates redundant recomputation in diffusion reinforcement learning by reusing the same feed-forward backbone for rollout and policy update, thereby avoiding unnecessary gradient tracking. It introduces two training schedules—LeanGRPO‑Retain, which reuses computation graphs and activations, and LeanGRPO‑Reweight, which backpropagates provisional gradients and corrects them later. These methods achieve up to a 1.83× speedup on FlowGRPO/DanceGRPO with FLUX.1‑dev and Wan while preserving the original optimization objective.