arXiv Machine Learning

Generalized Guarantees for Variational Inference in the Presence of Even and Elliptical Symmetry

arXiv:2511. 01064v3 Announce Type: replace-cross Abstract: Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions.

arXiv AI
Sep 18

Spherical Cauchy Variational Autoencoders: Heavy Angular Tails and Exact KL Evaluation

The paper introduces the spherical Cauchy distribution as a new hyperspherical posterior for variational autoencoders, avoiding the complications of the von Mises–Fisher and Power Spherical alternatives. By using stereographic projection and a Möbius transformation, the authors obtain exact posterior samples and a closed‑form KL divergence that terminates in a finite polynomial for even dimensions and admits certified truncation for odd dimensions. Empirical results show that the spherical Cauchy yields faster inference and lower reconstruction loss on MNIST and improved negative log‑likelihood on smallNORB compared to existing methods.

By Lukas Sablica, Kurt Hornik
arXiv Machine Learning
Sep 25

On the SoS Certifiability of Log-Concave Distributions

arXiv:2609. 30105v1 Announce Type: new Abstract: For an arbitrary isotropic log-concave distribution $P$ on $\mathbb{R}^d$, we prove that the polynomial $(Cm)^m\|v\|_2^m - \mathbb{E}_{X\sim P}\langle X,v\rangle^m$ is a sum of squares for every even $m\ge2$, where $C>0$ is a universal constant.

By Aleksandr Storozhenko
arXiv Machine Learning
Jul 28

Beyond ICA: Identifiability by Symmetry Breaking

arXiv:2607. 23182v1 Announce Type: cross Abstract: We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting.

By Pengzhou Wu
arXiv Machine Learning
Sep 15

Riemannian ascent--descent for nonconvex nonconcave minimax landscapes: convergence to basin saddle points and applications to distributionally robust optimization

The paper introduces a new convergence framework for solving distributionally robust optimization problems formulated as nonconvex, nonconcave minimax problems over a Euclidean space and a Riemannian manifold. It defines a "basin saddle point"—a locally defined Nash equilibrium—and proves that a Riemannian gradient ascent–descent algorithm converges to such points under a local Łojasiewicz growth condition. The authors apply this theory to a statistical risk DRO problem over Gaussian measures, deriving explicit convergence rates and constants in terms of data dimension, loss moments, and reference covariance.

By Rishabh Dixit, Pranav Upadrashta, Alex Cloninger
arXiv Machine Learning
Jun 2

Robust Learning of a Group DRO Neuron

arXiv:2601. 18115v2 Announce Type: replace Abstract: We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional shifts, for a broad family of covariate distributions.

By Guyang Cao, Shuyao Li, Sushrut Karmalkar, Jelena Diakonikolas
arXiv Statistics ML
Aug 27

Barycentric Weak Inner-Product Gromov-Wasserstein

The paper introduces a weak Gromov-Wasserstein (wGW) framework that compares source relations with relations between target conditional laws, focusing on inner-product relations and preserving conditional means. It defines the barycentric weak inner-product GW (wIGW) distance, proves existence of minimizers under finite second moments, and presents a ridge-regularized dual formulation leading to an iterative algorithm for finitely supported measures. Experiments on point clouds, graphs, and a PBMC multiome study demonstrate that mean-preserving target refinements can incur zero cost and improve atlas-based cell type transfer.

By Youssef Mroueh