arXiv Machine Learning

Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards

arXiv:2603. 23086v2 Announce Type: replace Abstract: Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and diversity.

arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv Computer Vision
Sep 7

Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

The paper introduces Reflection-Aware GRPO (RA‑GRPO), a reinforcement‑learning framework that aligns diffusion generative models with human preferences. It uses Diffusion Reflection to correct intermediate sampling paths by reversing the diffusion process, and Counterfactual Path Synthesis to embed these corrected trajectories into the policy, avoiding extra inference cost. Experiments on text‑to‑image and text‑to‑video models show RA‑GRPO outperforms existing methods, reducing reward hacking and improving generalization while remaining architecture‑agnostic.

By Junlong Wu, Jiuzhou Lin, Jia Sun, Boheng Zhang, Huaiqing Wang, Dewen Fan, Houde Liu, Qianqian Gan, Fan Yang, Tingting Gao
arXiv Computer Vision
Sep 7

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

The paper introduces Objective-aware Trajectory Credit Assignment (OTCA), a framework that refines reinforcement learning for diffusion-based visual generation. OTCA decomposes credit across denoising steps and allocates multiple reward signals adaptively, addressing the coarse, uniform reward assignment of existing GRPO pipelines. Experiments demonstrate that OTCA consistently enhances image and video generation quality across various metrics.

By Rui Li, Ke Hao, Yuanzhi Liang, Haibin Huang, Chi Zhang, Yun Gu, Xuelong Li
arXiv AI
Sep 18

Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design

The paper investigates how reinforcement learning can be effectively applied to diffusion models for visual tasks, focusing on the role of likelihood estimation. By systematically separating policy‑gradient objectives, likelihood estimators, and rollout sampling schemes, the authors find that using an evidence lower bound (ELBO) based likelihood estimator computed from the final generated sample is the key factor for stable and efficient RL optimization, outweighing the choice of loss function. Experiments on SD 3.5 Medium across multiple reward benchmarks confirm that this approach improves GenEval scores from 0.24 to 0.95 in 90 GPU hours, outperforming existing methods such as FlowGRPO and the current state‑of‑the‑art without reward hacking.

By Jaemoo Choi, Yuchen Zhu, Wei Guo, Petr Molodyk, Bo Yuan, Jinbin Bai, Yi Xin, Molei Tao, Yongxin Chen
arXiv AI
Sep 10

SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration

SwiftExplorer is a training‑free diffusion model alignment plugin that addresses two key issues in objective‑guided sampling: the loss of diversity due to strong directional bias and the inefficiency of constant guidance. It introduces an Inheritance‑Restart exploration mechanism to prevent early convergence and enhance the likelihood of high‑reward trajectories, while a Quality‑Efficiency arbitration mechanism removes incorrect signals and dynamically stops generation when optimal reward gain is achieved. Experiments show that SwiftExplorer improves preference, fidelity, diversity, and richness across multiple evaluation metrics.

By Renye Yan, Jikang Cheng, You Wu, Bojin Huang, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai
arXiv Computer Vision
Sep 25

AdaPilot: Towards Scene-Adaptive Policy Learning for Cross-Generator Text-to-Image Quality Optimization

AdaPilot introduces a scene-adaptive, cross-generator policy for optimizing text-to-image generation quality. By framing multi-turn image generation as a Markov Decision Process and using reinforcement learning, it decouples the policy from specific generator internals, incorporates scene-aware and process-level rewards, and achieves superior quality and generalization compared to baselines. Experiments demonstrate that a single AdaPilot policy can transfer zero‑shot to unseen generators while consistently improving performance across all evaluated models.

By Wenjin Liu, Fayuan Ke, Yue Lu, Zhe Cui, Anh Tuan Luu, Haoran Luo