arXiv Machine Learning

Rethinking Bregman Divergences in Kronecker-Factored Optimizers

arXiv:2606. 00542v1 Announce Type: new Abstract: Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures.

arXiv Machine Learning
Sep 16

Near-Optimal Nonconvex Matrix Completion

arXiv:2609. 17048v1 Announce Type: cross Abstract: We study nonconvex methods for matrix completion, the problem of recovering a low-rank matrix from a subset of its entries.

By Jian-Feng Cai, Xiliang Lu, Juntao You
Hugging Face Trending Papers
Sep 3

Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size

The paper introduces a Projected Riemannian Gradient Descent (RGD) algorithm for computing the Bures‑Wasserstein barycenter of positive definite matrices, achieving dimension‑independent linear convergence at unit step size. It resolves a previous dichotomy by showing that clipping eigenvalues to a fixed interval yields a closed‑form, non‑expansive projection in the BW metric, allowing the algorithm to match the empirical speed of unit‑step RGD while maintaining theoretical guarantees. The method also extends to the invariant matrix projection problem, providing a unified dimension‑independent analysis.

arXiv Machine Learning
Aug 20

Inference and Uncertainty Quantification for Streaming $r$-PCA

The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.

By Haoshu Xu, Hongzhe Li
arXiv AI
Sep 15

WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians

The paper introduces WaterKron, a method that integrates two-sided GPTQ with row- and column-dependent waterfilling scales and entropy coding for post‑training quantization. It derives a high‑rate distortion measure relative to the full Hessian, introducing a Kronecker‑Hessian mismatch factor Φ that quantifies the distortion penalty of using a Kronecker approximation. Minimizing Φ leads to a Gaussian covariance‑fitting problem solved via classical flip‑flop updates, yielding a FlipFlop Hessian that empirically improves KL divergence and perplexity compared to other Hessian choices.

By Johann Birnick, Rayan Saab