PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2607. 24583v1 Announce Type: new Abstract: Large scale Bayesian nonparametrics (BNP) learner such as Stochastic Variational Inference (SVI) can handle datasets with large class number and large training size at fractional cost.
The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.
Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density. However, this choice of latent density can impose unnecessary complexity on the learned flow transformation due to the topological mismatch between the latent and data densities, leading to slower training and suboptimal performance.
arXiv:2606. 29724v1 Announce Type: new Abstract: Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density.
arXiv:2509.02109v3 Announce Type: replace-cross Abstract: The Expectation-Maximisation (EM) algorithm is a central tool in statistics and machine learning, widely used for latent-variable models such...
The paper introduces a highly efficient variational approximation for Gaussian Mixture Models (GMMs) with arbitrary covariances, integrated with mixtures of factor analyzers. This method reduces the per‑iteration runtime from ≠O(NCD^2) to a complexity that scales linearly with dimensionality D and sublinearly with the product NC. Experiments demonstrate sublinear scaling across the entire optimization, order‑of‑magnitude speed‑ups on large benchmarks, training of GMMs with over 10 billion parameters in under nine hours on a single CPU, and competitive zero‑shot image denoising performance.