Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Pixel‑Space Diffusion via Observation Operators introduces a new framework for pixel‑space diffusion models that addresses a scale‑time mismatch in existing methods. By replacing fixed full‑image supervision with a time‑indexed observation trajectory that progresses from coarse structures to the full image, the model aligns supervision with the natural recovery order of image details. The approach employs Gaussian‑Lanczos operators and a GL‑CoDA decoder to refine features progressively, resulting in faster convergence and higher generation quality, achieving an FID of 1.52 on ImageNet‑256.
arXiv:2608.29322v1 Announce Type: new Abstract: Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or c...
The paper introduces a Posterior‑Dynamics Framework that leverages pretrained diffusion models as multiscale priors for linear imaging inverse problems such as deblurring, super‑resolution, and inpainting. By constructing a surrogate likelihood centered on the clean image and incorporating diffusion uncertainty, the authors derive continuous posterior dynamics and a tunable Langevin component for adaptive exploration. They prove theoretical guarantees (endpoint consistency, finite‑horizon tracking, weak accuracy) and present the PD‑IMEX sampler, which achieves high‑quality reconstructions with only 100 score evaluations and controllable fidelity‑diversity trade‑offs.
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
arXiv:2607. 09753v1 Announce Type: cross Abstract: Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that parameterize the score function.
The paper introduces CM-RED, a fast MRI reconstruction method that combines a pretrained consistency model with the regularization by denoising framework. By integrating controlled noise injection into accelerated proximal gradient updates, CM-RED achieves high‑quality reconstructions on fastMRI knee and brain datasets with only four network function evaluations. It consistently outperforms existing diffusion‑ and consistency‑based approaches in quantitative metrics, visual fidelity, and robustness to hyperparameter changes.