The paper introduces ZeNOVA, a gradient‑free method for aligning initial noise in generative models. It uses annealed soft‑value guidance, manifold‑constrained hyperspherical Langevin dynamics, and Metropolis‑Hastings jumps to address instability in black‑box reward settings. Experiments on image and video models show ZeNOVA outperforms existing zeroth‑order baselines by more stably optimizing noise toward higher rewards.
By Jinho Chang, Jong Chul Ye
arXiv:2511.21415v2 Announce Type: replace
Abstract: We introduce DiverseVAR, a framework that enhances the diversity of text-conditioned visual autoregressive models (VAR) at test time without requir...
By Mingue Park, Prin Phunyaphibarn, Phillip Y. Lee, Minhyuk Sung
CrossDistill is a trajectory-level hybrid few-step distillation framework for diffusion models that balances quality and diversity by splitting the sampling trajectory at a crossover point. The high-noise interval uses a trajectory-preserving objective to maintain global mode coverage, while the low-noise interval applies a distribution-matching objective to sharpen local details, with the two stages coupled through the crossover state. This noise-level scheduling policy, demonstrated on text-to-video and image-to-video diffusion models, expands the few-step quality-diversity frontier by preserving seed-level variation while achieving competitive visual fidelity.
By Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang
arXiv:2603. 28762v2 Announce Type: replace-cross Abstract: Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt.
By Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or
The paper introduces a Posterior‑Dynamics Framework that leverages pretrained diffusion models as multiscale priors for linear imaging inverse problems such as deblurring, super‑resolution, and inpainting. By constructing a surrogate likelihood centered on the clean image and incorporating diffusion uncertainty, the authors derive continuous posterior dynamics and a tunable Langevin component for adaptive exploration. They prove theoretical guarantees (endpoint consistency, finite‑horizon tracking, weak accuracy) and present the PD‑IMEX sampler, which achieves high‑quality reconstructions with only 100 score evaluations and controllable fidelity‑diversity trade‑offs.
By Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng
arXiv:2608.29335v1 Announce Type: new
Abstract: Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on th...
By Guangting Zheng, Yiyuan Zhang, Tao Yang, Yunpeng Chen, Rui Zhu, Jiajun Deng, Yanyong Zhang