arXiv Machine Learning

Symmetrizing Bregman Divergence on the Cone of Positive Definite Matrices: Which Mean to Use and Why

arXiv Machine Learning
Sep 14

Information-Induced Training Geometry: Exact Reduction, Canonical Completion, and Structured Expressivity

The paper investigates how training data limits the geometry of an optimizer through the covectors that a specified information channel can observe. It establishes that, under affine‑invariant Riemannian geometry, a full‑column‑rank positive‑definite compression can be uniquely completed via a split‑Hadamard metric submetry, yielding exact variational reduction from the full geometry to the visible target. The resulting framework provides explicit formulas for pullback metrics, Gram matrices, and prior‑data shrinkage, and characterizes the gauge‑invariant rank stratification of the positive‑definite cone as the channel varies.

By Zavier Li
arXiv Machine Learning
Sep 14

Block-Norm Geometries for Online Mirror Descent with Sparse Losses

The paper investigates how the choice of geometry in online mirror descent affects performance, particularly when loss gradients are sparse. It introduces randomized block‑norm mirror maps that interpolate between Euclidean and entropic geometries, achieving polynomial‑in‑dimension regret improvements over standard methods for various convex sets. The authors also demonstrate that naive alternation between mirror maps can lead to linear regret and propose a Hedge‑based meta‑algorithm that competes with the best mirror map in a finite portfolio, achieving near‑optimal regret for random block geometries.

By Swati Gupta, Jai Moondra, Mohit Singh
arXiv Machine Learning
Jun 16

The Reverse Telescoping Coordinate System for Positive Definite Matrices: Geometry, Computation, and Generative Modeling

arXiv:2606. 15442v1 Announce Type: cross Abstract: We design a new unconstrained coordinate system where a $p\times p$ symmetric positive definite (SPD) matrix $\Theta$ is represented by a reverse telescoping map $\Theta(x)=\rm{RT}(x)$, with $x=(v,d,r)\in\mathbb{R}\times\mathbb{R}^{(p-1)}\times\mathbb{R}^{p(p-1)/2}$, representing respectively the log volume or log determinant; and the shape, as encoded by log relative diagonal scales and partial covariances among the nodes.

By Anindya Bhadra
arXiv Machine Learning
Jul 13

Group Invariant Spectral Embedding

arXiv:2607. 08987v1 Announce Type: new Abstract: Spectral embedding methods are widely used for dimensionality reduction and clustering of high-dimensional datasets with intrinsic low-dimensional structures.

By Yeari Vigder, Paulina Hoyos, David Thong, Joakim and\'en, Joe Kileel, Amit Moscovich
arXiv Machine Learning
Jul 28

Beyond ICA: Identifiability by Symmetry Breaking

arXiv:2607. 23182v1 Announce Type: cross Abstract: We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting.

By Pengzhou Wu
arXiv Machine Learning
Jun 2

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

arXiv:2606. 01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling is challenging due to the presence of additional symmetries under coupled row/column scalings between the two factors.

By Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra