Symmetrizing Bregman Divergence on the Cone of Positive Definite Matrices: Which Mean to Use and Why
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2511. 01064v3 Announce Type: replace-cross Abstract: Variational inference (VI) approximates a target density $p$ by the best match $q$ in a family of tractable distributions.
The paper investigates how training data limits the geometry of an optimizer through the covectors that a specified information channel can observe. It establishes that, under affine‑invariant Riemannian geometry, a full‑column‑rank positive‑definite compression can be uniquely completed via a split‑Hadamard metric submetry, yielding exact variational reduction from the full geometry to the visible target. The resulting framework provides explicit formulas for pullback metrics, Gram matrices, and prior‑data shrinkage, and characterizes the gauge‑invariant rank stratification of the positive‑definite cone as the channel varies.
The paper investigates how the choice of geometry in online mirror descent affects performance, particularly when loss gradients are sparse. It introduces randomized block‑norm mirror maps that interpolate between Euclidean and entropic geometries, achieving polynomial‑in‑dimension regret improvements over standard methods for various convex sets. The authors also demonstrate that naive alternation between mirror maps can lead to linear regret and propose a Hedge‑based meta‑algorithm that competes with the best mirror map in a finite portfolio, achieving near‑optimal regret for random block geometries.
arXiv:2606. 15442v1 Announce Type: cross Abstract: We design a new unconstrained coordinate system where a $p\times p$ symmetric positive definite (SPD) matrix $\Theta$ is represented by a reverse telescoping map $\Theta(x)=\rm{RT}(x)$, with $x=(v,d,r)\in\mathbb{R}\times\mathbb{R}^{(p-1)}\times\mathbb{R}^{p(p-1)/2}$, representing respectively the log volume or log determinant; and the shape, as encoded by log relative diagonal scales and partial covariances among the nodes.
arXiv:2607. 08987v1 Announce Type: new Abstract: Spectral embedding methods are widely used for dimensionality reduction and clustering of high-dimensional datasets with intrinsic low-dimensional structures.
arXiv:2502. 00753v4 Announce Type: replace-cross Abstract: Smoothness is crucial for attaining fast rates in first-order optimization.