arXiv Computer Vision By Rui Xia, Ayan Das, Artem Artemev, Andi Zhang, Guillaume Hennequin, Alberto Bernacchia

Improved denoising diffusion probabilistic models with efficient non-diagonal covariance modeling

Read the original on arXiv Computer Vision →

The paper proposes a new covariance model for Denoising Diffusion Probabilistic Models (DDPMs) that captures non‑diagonal correlations and the power‑law frequency spectrum of natural images. Using a Kronecker‑factored DCT (K‑DCT) decomposition, the authors reduce computational complexity from quadratic to log‑linear, enabling efficient sampling with few steps. Experiments on CIFAR‑10, Celeb‑A, ImageNet, and LSUN demonstrate improved FID and likelihoods over previous state‑of‑the‑art samplers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Sep 11

Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters

The paper introduces a highly efficient variational approximation for Gaussian Mixture Models (GMMs) with arbitrary covariances, integrated with mixtures of factor analyzers. This method reduces the per‑iteration runtime from ≠O(NCD^2) to a complexity that scales linearly with dimensionality D and sublinearly with the product NC. Experiments demonstrate sublinear scaling across the entire optimization, order‑of‑magnitude speed‑ups on large benchmarks, training of GMMs with over 10 billion parameters in under nine hours on a single CPU, and competitive zero‑shot image denoising performance.

By Sebastian Salwig, Till Kahlke, Florian Hirschberger, Dennis Forster, J\"org L\"ucke
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal