arXiv AI

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.

arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv AI
Sep 10

SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration

SwiftExplorer is a training‑free diffusion model alignment plugin that addresses two key issues in objective‑guided sampling: the loss of diversity due to strong directional bias and the inefficiency of constant guidance. It introduces an Inheritance‑Restart exploration mechanism to prevent early convergence and enhance the likelihood of high‑reward trajectories, while a Quality‑Efficiency arbitration mechanism removes incorrect signals and dynamically stops generation when optimal reward gain is achieved. Experiments show that SwiftExplorer improves preference, fidelity, diversity, and richness across multiple evaluation metrics.

By Renye Yan, Jikang Cheng, You Wu, Bojin Huang, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai
Hugging Face Trending Papers
Aug 10

Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation

Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level.

arXiv Machine Learning
Jun 16

DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning

arXiv:2505. 09655v5 Announce Type: replace-cross Abstract: Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning.

By Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi
arXiv Machine Learning
1d ago

Continual Reinforcement Learning with Neuroevolution

The paper investigates continual reinforcement learning using neuroevolution, comparing evolution strategies (ES) and genetic algorithms (GAs) across diverse environments and network sizes. ES consistently achieves a better balance between stability and plasticity, while GAs are more plastic but forget more. The authors attribute this to ES finding wider neighborhoods in weight space, with overlap between consecutive tasks correlating with the stability-plasticity trade‑off, and note that common RL plasticity issues do not transfer to neuroevolution.

By Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi
arXiv AI
3d ago

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

Fenchel Tilt Flow Control (FTFC) is a new method for fine‑tuning pretrained generative models to arbitrary preference functions. It decouples utility optimization from model fitting by first learning reward and density‑ratio weights on pretrained samples, then freezing these weights to adjust a diffusion or flow model in a single importance‑weighted stage. The approach supports general f‑divergence penalties, achieves exact duality for concave utilities, and demonstrates up to 20× efficiency gains while outperforming baselines on image and molecule generation tasks.

By Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov
arXiv AI
Sep 1

RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

RAGDiffusion++ advances garment generation by addressing the high‑frequency texture gap that previous retrieval‑augmented models left unresolved. The approach introduces a dual‑image FLUX architecture trained on a large, complex garment dataset, coupled with a new attribute‑aware reward model that guides reinforcement learning to favor realistic high‑frequency patterns. An adversarial‑regularized RL strategy (AR‑GRPO) further prevents artifact exploitation, ensuring the model samples authentic, detailed garment textures.

By Yuhan Li, Xianfeng Tan, Fangao Zeng, Wenxiang Shang, Pipei Huang, Hao Zhou, Zhiyu Jin, Wenjun Zhang, Bingbing Ni