arXiv AI By Renye Yan, Jikang Cheng, You Wu, Bojin Huang, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai

SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration

Read the original on arXiv AI →

SwiftExplorer is a training‑free diffusion model alignment plugin that addresses two key issues in objective‑guided sampling: the loss of diversity due to strong directional bias and the inefficiency of constant guidance. It introduces an Inheritance‑Restart exploration mechanism to prevent early convergence and enhance the likelihood of high‑reward trajectories, while a Quality‑Efficiency arbitration mechanism removes incorrect signals and dynamically stops generation when optimal reward gain is achieved. Experiments show that SwiftExplorer improves preference, fidelity, diversity, and richness across multiple evaluation metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 6

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.

By Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
arXiv AI
2d ago

Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

COFFEE is a plug‑and‑play framework that enables future‑aware guidance for discrete diffusion models by separating sequence dependence from the objective. It uses a target‑free carrier to absorb marginal token distributions and a compiled finite‑state model to capture how token combinations affect sequence‑level preferences, allowing global preferences to be transferred to unresolved positions without retraining the diffusion model. The framework supports both hard constraints and learned soft objectives and demonstrates strong control results across symbolic, language, and biological benchmarks.

By Hua (Edward), Xu, Dongxin Li, Gwen Yidou-Weng, Guy Van den Broeck, Wei Wang, Anji Liu
arXiv Machine Learning
Aug 28

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

The paper introduces GRAS, a method that improves training‑free reward alignment for discrete diffusion models by reducing variance in guided proposals and adapting the resampling temperature during search. It achieves this without adding denoiser cost, using Rao‑Blackwellized estimates for differentiable rewards and a leave‑one‑out baseline for non‑differentiable ones. Experiments on regulatory DNA and protein design show GRAS outperforms existing training‑free techniques and rivals reward‑fine‑tuned models.

By Kwanyoung Kim
arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang