arXiv Machine Learning By Yuan Cao, Yifu Tang, Hangqi Li, Zeyu Zheng

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

Read the original on arXiv Machine Learning →

The paper studies how to allocate a fixed computational budget across the denoising steps of diffusion models to improve sample quality at deployment. It shows that the expected benefit of evaluating multiple candidates at a step can be decomposed into a step‑specific sensitivity and a universal sample‑size factor, and that the optimal allocation follows a water‑filling structure. Experiments demonstrate that this allocation achieves the same quality as a uniform strategy while reducing function evaluations by 20–50%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

Schedule optimization for tau-leaping in masked discrete diffusion

The paper studies how to choose sampling schedules for tau‑leaping in masked discrete diffusion models. By deriving an exact integral representation of the factorization error ε_fact in terms of a dependence density ρ, the authors develop estimators and recursive equations that identify the unique optimal schedule under a monotonicity condition. In the large‑scale limit, they provide explicit characterizations of the optimal smooth schedule and show that while optimizing smooth schedules can improve constants, it does not change the N/K scaling unless the dependence density degenerates, in which case asymptotic improvements are possible.

By Cecilia Secchi, Giacomo Zanella
arXiv Machine Learning
6d ago

Spectral-Guided Diffusion: Accelerating Inference via Static Spectral Layer Scheduling

Spectral-Guided Diffusion introduces a method to accelerate diffusion inference by identifying and reusing residual branches that need not be recomputed during the trajectory. The approach uses a Spectral Concentration Ratio (SCR) combined with Frobenius magnitude to create an offline sensitivity proxy and deterministic lifetime for each scheduled unit, eliminating the need for routers or input-dependent searches. Experiments on models such as LLaDA-8B, DiT-XL/2, U-ViT-L, and SDXL show that this scheduling preserves quality better than several baselines and achieves up to a 3.0× wall‑clock speedup over eager inference.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma