arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.
By Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
arXiv:2603. 23086v2 Announce Type: replace Abstract: Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and diversity.
By Orhun Bugra Baran, Melih Kandemir, Ramazan Gokberk Cinbis
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences.
COFFEE is a plug‑and‑play framework that enables future‑aware guidance for discrete diffusion models by separating sequence dependence from the objective. It uses a target‑free carrier to absorb marginal token distributions and a compiled finite‑state model to capture how token combinations affect sequence‑level preferences, allowing global preferences to be transferred to unresolved positions without retraining the diffusion model. The framework supports both hard constraints and learned soft objectives and demonstrates strong control results across symbolic, language, and biological benchmarks.
By Hua (Edward), Xu, Dongxin Li, Gwen Yidou-Weng, Guy Van den Broeck, Wei Wang, Anji Liu
The paper introduces GRAS, a method that improves training‑free reward alignment for discrete diffusion models by reducing variance in guided proposals and adapting the resampling temperature during search. It achieves this without adding denoiser cost, using Rao‑Blackwellized estimates for differentiable rewards and a leave‑one‑out baseline for non‑differentiable ones. Experiments on regulatory DNA and protein design show GRAS outperforms existing training‑free techniques and rivals reward‑fine‑tuned models.
By Kwanyoung Kim
arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.
By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv:2605. 30789v2 Announce Type: replace-cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs.
By Yiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi, Yukang Chen, Dingdong Wang, Tianhe Wu, Junjie Wang, Yujiu Yang, Yu Qiao, Ruihang Chu
arXiv:2608. 09385v1 Announce Type: cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself.
By Hossein Goli, Farzan Farnia, Amin Gohari
arXiv:2603. 14504v2 Announce Type: replace-cross Abstract: Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at inference time.
By Niklas Schweiger, Daniel Cremers, Karnik Ram
arXiv:2607. 00691v1 Announce Type: new Abstract: Black-box optimization is a fundamental science and engineering tool that makes it possible to optimize objectives without gradient information.
By Edouard R. Dufour, Pascal Fua
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level.
arXiv:2511.19811v2 Announce Type: replace-cross
Abstract: Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive output...
By Debin Meng, Chen Jin, Zheng Gao, Yanran Li, Ioannis Patras, Georgios Tzimiropoulos