arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.
By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
SwiftExplorer is a training‑free diffusion model alignment plugin that addresses two key issues in objective‑guided sampling: the loss of diversity due to strong directional bias and the inefficiency of constant guidance. It introduces an Inheritance‑Restart exploration mechanism to prevent early convergence and enhance the likelihood of high‑reward trajectories, while a Quality‑Efficiency arbitration mechanism removes incorrect signals and dynamically stops generation when optimal reward gain is achieved. Experiments show that SwiftExplorer improves preference, fidelity, diversity, and richness across multiple evaluation metrics.
By Renye Yan, Jikang Cheng, You Wu, Bojin Huang, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai
arXiv:2608. 09385v1 Announce Type: cross Abstract: Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself.
By Hossein Goli, Farzan Farnia, Amin Gohari
arXiv:2609.14896v1 Announce Type: cross
Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at infer...
By Jiayi Yuan, Hangoo Kang, James Jihao Liu, Yejin Choi, Vikram Iyer, Liwei Jiang, Natasha Jaques
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level.
arXiv:2606. 03962v1 Announce Type: cross Abstract: Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward.
By Anthony GX-Chen, Ankit Anand, Gheorghe Comanici, Zaheer Abbas, Eser Ayg\"un, David Smalling, Shibl Mourad, Doina Precup, Andr\'e Barreto, Mark Rowland
Diffusion models have strong generative capabilities. However, their maximum likelihood training objective only focuses on reconstructing the data distribution, making it difficult to align with specific preferences.
arXiv:2505. 09655v5 Announce Type: replace-cross Abstract: Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning.
By Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi
The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.
By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
Fenchel Tilt Flow Control (FTFC) is a new method for fine‑tuning pretrained generative models to arbitrary preference functions. It decouples utility optimization from model fitting by first learning reward and density‑ratio weights on pretrained samples, then freezing these weights to adjust a diffusion or flow model in a single importance‑weighted stage. The approach supports general f‑divergence penalties, achieves exact duality for concave utilities, and demonstrates up to 20× efficiency gains while outperforming baselines on image and molecule generation tasks.
By Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov
RAGDiffusion++ advances garment generation by addressing the high‑frequency texture gap that previous retrieval‑augmented models left unresolved. The approach introduces a dual‑image FLUX architecture trained on a large, complex garment dataset, coupled with a new attribute‑aware reward model that guides reinforcement learning to favor realistic high‑frequency patterns. An adversarial‑regularized RL strategy (AR‑GRPO) further prevents artifact exploitation, ensuring the model samples authentic, detailed garment textures.
By Yuhan Li, Xianfeng Tan, Fangao Zeng, Wenxiang Shang, Pipei Huang, Hao Zhou, Zhiyu Jin, Wenjun Zhang, Bingbing Ni
arXiv:2609.30840v1 Announce Type: cross
Abstract: One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit ge...
By Austin Wang, Ziheng Cheng, Lexing Ying