arXiv AI By Xiang Li, Dianbo Liu, Kenji Kawaguchi

Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

Read the original on arXiv AI →

arXiv:2606. 02453v1 Announce Type: cross Abstract: Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Manifold-Constrained Initial Noise Optimization for Efficient Generative Model Alignment

The paper introduces ZeNOVA, a gradient‑free method for aligning initial noise in generative models. It uses annealed soft‑value guidance, manifold‑constrained hyperspherical Langevin dynamics, and Metropolis‑Hastings jumps to address instability in black‑box reward settings. Experiments on image and video models show ZeNOVA outperforms existing zeroth‑order baselines by more stably optimizing noise toward higher rewards.

By Jinho Chang, Jong Chul Ye
arXiv Computer Vision
Sep 15

CrossDistill: Balancing Quality and Diversity via Trajectory-Level Hybrid Few-Step Distillation

CrossDistill is a trajectory-level hybrid few-step distillation framework for diffusion models that balances quality and diversity by splitting the sampling trajectory at a crossover point. The high-noise interval uses a trajectory-preserving objective to maintain global mode coverage, while the low-noise interval applies a distribution-matching objective to sharpen local details, with the two stages coupled through the crossover state. This noise-level scheduling policy, demonstrated on text-to-video and image-to-video diffusion models, expands the few-step quality-diversity frontier by preserving seed-level variation while achieving competitive visual fidelity.

By Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang
arXiv AI
Sep 4

A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors

The paper introduces a Posterior‑Dynamics Framework that leverages pretrained diffusion models as multiscale priors for linear imaging inverse problems such as deblurring, super‑resolution, and inpainting. By constructing a surrogate likelihood centered on the clean image and incorporating diffusion uncertainty, the authors derive continuous posterior dynamics and a tunable Langevin component for adaptive exploration. They prove theoretical guarantees (endpoint consistency, finite‑horizon tracking, weak accuracy) and present the PD‑IMEX sampler, which achieves high‑quality reconstructions with only 100 score evaluations and controllable fidelity‑diversity trade‑offs.

By Zhaoqiang Liu, Tongyao Pang, Ruibing Wang, Yang Zheng