arXiv Machine Learning

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion proposes a new inference-time framework that improves steering of frozen masked discrete diffusion models. By reducing the variance of guidance estimates, applying reward tilting to clean-token logits, and adapting the selection temperature at each step, VGAS addresses three default choices in existing pipelines. Experiments on regulatory DNA, protein, and small-molecule benchmarks show that VGAS achieves the best training-free reward performance and matches or surpasses reward-fine-tuned generators.

arXiv Machine Learning
Aug 28

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

The paper introduces GRAS, a method that improves training‑free reward alignment for discrete diffusion models by reducing variance in guided proposals and adapting the resampling temperature during search. It achieves this without adding denoiser cost, using Rao‑Blackwellized estimates for differentiable rewards and a leave‑one‑out baseline for non‑differentiable ones. Experiments on regulatory DNA and protein design show GRAS outperforms existing training‑free techniques and rivals reward‑fine‑tuned models.

By Kwanyoung Kim
arXiv Machine Learning
Jul 7

On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

arXiv:2607. 02834v1 Announce Type: new Abstract: Molecular optimization often starts from a pretrained generative model that captures a broad prior over valid molecular structures.

By Trevor Chen, Ariel Dai, Jason Yang, Riccardo De Santi, Daniel Khalil, Wenda Chu, Nate Gruver, Pranav Murugan, Alexander F. G. Goldberg, Maruan Al-Shedivat, Yisong Yue
arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv AI
4d ago

Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

Spectral Feedback is a new algorithm for aligning discrete diffusion models at test time by iteratively revisiting and editing token positions rather than only steering the reverse process. It selects edit-sets—groups of token positions to re-mask and re-sample—using sparse Fourier representations of edit-set value functions, enabling efficient optimization of which tokens to revisit. The method is model-agnostic and improves alignment performance across pretrained, test‑time aligned, and fine‑tuned diffusion models, achieving significant gains in protein stability for inverse folding tasks.

By Shai Dickman, Mert Cemri, Landon Butler, Kannan Ramchandran
arXiv Machine Learning
Jun 2

Drifting Preference Optimization for One-Step Generative Models

arXiv:2606. 02521v1 Announce Type: new Abstract: One-step text-to-image generators are attractive for deployment because they generate an image with a single forward pass, but preference finetuning them remains difficult: standard alignment methods often rely on policy likelihoods, denoising trajectories, differentiable reward gradients, or test-time optimization.

By Zhou Jiang, Yandong Wen, Zhen Liu
arXiv AI
1d ago

Fenchel Tilting: Weighted Correction for Efficient Finetuning of Generative Models

Fenchel Tilt Flow Control (FTFC) is a new method for fine‑tuning pretrained generative models to arbitrary preference functions. It decouples utility optimization from model fitting by first learning reward and density‑ratio weights on pretrained samples, then freezing these weights to adjust a diffusion or flow model in a single importance‑weighted stage. The approach supports general f‑divergence penalties, achieves exact duality for concave utilities, and demonstrates up to 20× efficiency gains while outperforming baselines on image and molecule generation tasks.

By Maksim Bobrin, Maksim Zhdanov, Dmitry Dylov