arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.
By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv:2602. 08646v3 Announce Type: replace Abstract: We propose a gradient preconditioning method that makes reward-guided generation with one-step generative models both efficient and reliable.
By Jisung Hwang, Minhyuk Sung
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences.
arXiv:2603. 14504v2 Announce Type: replace-cross Abstract: Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at inference time.
By Niklas Schweiger, Daniel Cremers, Karnik Ram
The paper introduces GRAS, a method that improves training‑free reward alignment for discrete diffusion models by reducing variance in guided proposals and adapting the resampling temperature during search. It achieves this without adding denoiser cost, using Rao‑Blackwellized estimates for differentiable rewards and a leave‑one‑out baseline for non‑differentiable ones. Experiments on regulatory DNA and protein design show GRAS outperforms existing training‑free techniques and rivals reward‑fine‑tuned models.
By Kwanyoung Kim
The paper introduces Diffusion LAIR, a listwise preference optimization technique that leverages continuous reward scores instead of binary pairwise comparisons to align text‑to‑image diffusion models. LAIR transforms reward scores into centered advantage weights and optimizes an advantage‑weighted regression objective on an implicit reward defined by denoising‑loss improvement over a reference model, with a quadratic penalty to regulate reward magnitude. Experiments demonstrate that Diffusion LAIR surpasses strong baseline methods on SD1.5 and SDXL across generation, compositional, and editing tasks.
By Austin Wang, Jiaqi Han, Stefano Ermon, Yisong Yue
arXiv:2604. 17415v3 Announce Type: replace-cross Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model.
By Jeongjae Lee, Jinho Chang, Jeongsol Kim, Jong Chul Ye
arXiv:2609.33803v2 Announce Type: replace-cross
Abstract: Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate...
By Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wang, Ziqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu
arXiv:2609.38329v1 Announce Type: new
Abstract: Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This...
By Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
arXiv:2607. 07693v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.
By Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We in...
The paper introduces ZeNOVA, a gradient‑free method for aligning initial noise in generative models. It uses annealed soft‑value guidance, manifold‑constrained hyperspherical Langevin dynamics, and Metropolis‑Hastings jumps to address instability in black‑box reward settings. Experiments on image and video models show ZeNOVA outperforms existing zeroth‑order baselines by more stably optimizing noise toward higher rewards.
By Jinho Chang, Jong Chul Ye