arXiv Computer Vision

Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

The paper proposes a new covariance model for Denoising Diffusion Probabilistic Models (DDPMs) that captures non‑diagonal correlations and the power‑law frequency spectrum of natural images. Using a Kronecker‑factored DCT (K‑DCT) decomposition, the authors reduce computational complexity from quadratic to log‑linear, enabling efficient sampling with few steps. Experiments on CIFAR‑10, Celeb‑A, ImageNet, and LSUN demonstrate improved FID and likelihoods over previous state‑of‑the‑art samplers.

arXiv Machine Learning
Sep 11

Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters

The paper introduces a highly efficient variational approximation for Gaussian Mixture Models (GMMs) with arbitrary covariances, integrated with mixtures of factor analyzers. This method reduces the per‑iteration runtime from ≠O(NCD^2) to a complexity that scales linearly with dimensionality D and sublinearly with the product NC. Experiments demonstrate sublinear scaling across the entire optimization, order‑of‑magnitude speed‑ups on large benchmarks, training of GMMs with over 10 billion parameters in under nine hours on a single CPU, and competitive zero‑shot image denoising performance.

By Sebastian Salwig, Till Kahlke, Florian Hirschberger, Dennis Forster, J\"org L\"ucke
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv AI
Sep 25

FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference

FB‑GDM is a fully‑Bayesian guided diffusion method that eliminates the need for task‑specific hyperparameter tuning in linear inverse problems. It derives a closed‑form conditional score from a Gaussian approximation of ΦGDM, treating two precision parameters as latent variables inferred via variational inference at each reverse step. Experiments on CelebA‑HQ show that FB‑GDM outperforms ΦGDM at its nominal setting, matches a ground‑truth‑calibrated oracle within 0.1 dB, and remains robust to changes in the forward operator, noise level, or image distribution without hallucinations.

By Gatien S\'eguy (SATIE), Thomas Rodet (SATIE)
arXiv Computer Vision
Sep 24

ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming

ZoomDiff is a high‑fidelity diffusion model designed to improve dual‑camera smooth zooming by producing photo‑realistic transitions. It strengthens dual‑image conditional guidance during denoising, injects flow‑aligned multi‑scale features from the VAE encoder into the decoder to recover high‑frequency details, and uses flow‑guided temporal consistency supervision to ensure smoother transitions. Experiments on synthetic and real‑world datasets show that ZoomDiff outperforms state‑of‑the‑art methods both quantitatively and qualitatively.

By Jiayi Zhang, Renlong Wu, Yukang Ding, Sibin Deng, Wangmeng Zuo