arXiv AI By Dake Bu, Wei Huang, Andi Han, Si Wu, Hau-San Wong, Qingfu Zhang, Taiji Suzuki, Atsushi Nitanda

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models

Read the original on arXiv AI →

The paper introduces DPRM, a plug‑in token‑ordering module for discrete diffusion models that uses a Doob h‑transform to convert terminal rewards into per‑position process rewards. By estimating these rewards from generation progress, confidence, and optional state, DPRM reorders tokens without altering the underlying model or sampler. Experiments on nine open‑source hosts show significant gains in reasoning, numeric VQA, visual‑codebook ordering, and preference‑conditioned generation, with improvements ranging from 8.97 to 53.3 points across diverse tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Diffusion Reward Models

arXiv:2609.33803v2 Announce Type: replace-cross Abstract: Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate...

By Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wang, Ziqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu
arXiv Machine Learning
Sep 18

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion proposes a new inference-time framework that improves steering of frozen masked discrete diffusion models. By reducing the variance of guidance estimates, applying reward tilting to clean-token logits, and adapting the selection temperature at each step, VGAS addresses three default choices in existing pipelines. Experiments on regulatory DNA, protein, and small-molecule benchmarks show that VGAS achieves the best training-free reward performance and matches or surpasses reward-fine-tuned generators.

By Kwanyoung Kim
arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun