arXiv Machine Learning

E$^2$M: Double Bounded $\alpha$-Divergence Optimization for Tensor-based Discrete Density Estimation

arXiv:2405. 18220v4 Announce Type: replace-cross Abstract: Tensor-based discrete density estimation requires flexible modeling and proper divergence criteria to enable effective learning; however, traditional approaches using $\alpha$-divergence face analytical challenges due to the $\alpha$-power terms in the objective function, which hinder the derivation of closed-form update rules.

arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv Machine Learning
Jun 4

Low-rank Distributional Matrix Completion

arXiv:2606. 04176v1 Announce Type: new Abstract: We study a distributional generalization of the matrix completion problem in which each entry of the target matrix is a probability distribution rather than a scalar.

By Jiayi Wang, Raymond K. W. Wong
arXiv Machine Learning
23h ago

Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

arXiv:2506. 06584v2 Announce Type: replace Abstract: Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorithm and its popular variant gradient EM being arguably the most widely used algorithms in practice.

By Mo Zhou, Weihang Xu, Maryam Fazel, Simon S. Du