OpenAI Blog

PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications

arXiv Statistics ML
Sep 25

Riemannian Gradient Descent for Gaussian Mixture Models with unknown diagonal covariances

The paper studies the numerical solution of the Beurling‑LASSO (BLASSO) for estimating Gaussian mixture models (GMMs) with unknown numbers of components and unknown diagonal covariance matrices. It introduces a Conic Particle Gradient Descent (CPGD) algorithm that incorporates Riemannian gradient descent to respect the Fisher‑Rao geometry of Gaussian distributions. The authors provide theoretical convergence guarantees, including exponential local convergence under a non‑degeneracy condition related to component separation, and demonstrate through numerical experiments that CPGD is more robust to overspecification of components than the EM algorithm.

By Romane Giard, Yohann De Castro, Roland Denis, Cl\'ement Marteau
Hugging Face Trending Papers
Jun 29

Simplifying Flow Matching Transformations with Low-Rank Mixture Models

Normalizing flows are powerful generative models that learn an invertible mapping between complex data distributions and simple latent distributions, typically a standard normal density. However, this choice of latent density can impose unnecessary complexity on the learned flow transformation due to the topological mismatch between the latent and data densities, leading to slower training and suboptimal performance.

arXiv Machine Learning
Sep 11

Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters

The paper introduces a highly efficient variational approximation for Gaussian Mixture Models (GMMs) with arbitrary covariances, integrated with mixtures of factor analyzers. This method reduces the per‑iteration runtime from ≠O(NCD^2) to a complexity that scales linearly with dimensionality D and sublinearly with the product NC. Experiments demonstrate sublinear scaling across the entire optimization, order‑of‑magnitude speed‑ups on large benchmarks, training of GMMs with over 10 billion parameters in under nine hours on a single CPU, and competitive zero‑shot image denoising performance.

By Sebastian Salwig, Till Kahlke, Florian Hirschberger, Dennis Forster, J\"org L\"ucke
arXiv Machine Learning
Aug 24

Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score

The paper introduces an amortized learning framework for selecting bandwidths in kernel density estimation by optimizing the logarithmic score across a distribution of tasks. It uses a truncated-and-renormalized bounded-support formulation and affine standardization to achieve stable learning and transferability across different intervals. Experiments on Gaussian samples, a multi-family benchmark, and randomized Gaussian mixtures demonstrate that the learned selector outperforms traditional methods such as Silverman’s rule, Sheather–Jones, and least‑squares cross‑validation, especially for small or heterogeneous samples.

By Junyi Liang, Hailiang Du
arXiv Machine Learning
Sep 23

Learning to bin: differentiable and Bayesian optimization for multi-dimensional discriminants in high-energy physics

The paper introduces a method for optimizing bin boundaries in multi-dimensional discriminants using a Gaussian Mixture Model (GMM), allowing flexible definition of analysis categories. Two optimization strategies—differentiable and Bayesian—are compared in toy binary and three-class setups, with the differentiable approach excelling in multi-dimensional cases. Applied to the FAIR Universe $H ightarrow au au$ dataset, the GMM-based optimization achieves the highest signal significance, and the tools are released as lightweight Python plugins.

By Johannes Erdmann, Nitish Kumar Kasaraguppe, Florian Mausolf
arXiv AI
Jun 16

FastMix: Fast Data Mixture Optimization via Gradient Descent

arXiv:2606. 14971v1 Announce Type: cross Abstract: While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and post-training remains a significant open problem.

By Haoru Tan, Sitong Wu, Yanfeng Chen, Jun Xia, Ruobing Xie, Bin Xia, Xingwu Sun, Xiaojuan Qi