arXiv Statistics ML

An efficient EM algorithm for both element-wise and structural missingness in matrix-variate normal mixture models

arXiv Machine Learning
5d ago

Robust Graph Clustering Network for Multiple Missing Data

The paper introduces the Robust Graph Clustering Network for Multiple Missing Data (RGCN), a method designed to cluster graphs with simultaneous missing node attributes and structural links. RGCN employs a view‑decoupled dual‑branch imputation to reduce cross‑view interference, a multi‑hyperspherical mixture prior to improve cluster compactness and separability on a directional latent manifold, and a boundary‑aware contrastive enhancement objective to counteract cluster blurring caused by imputation bias. Experiments on real‑world datasets show that RGCN consistently outperforms state‑of‑the‑art baselines across various missing data patterns.

By Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
arXiv Statistics ML
Sep 22

Tensor Completion using Subspace Information

Tensor Completion using Subspace Information (TCSI) is an algorithm that leverages side information by estimating a subspace and reformulating tensor completion as a matrix regression problem. Theoretical analysis shows that accurate subspace information reduces sample complexity to nearly linear in the uncoupled ambient dimensions and relaxes signal-to-noise ratio requirements compared to existing guarantees. Numerical simulations and an application to reconstructing global Total Electron Content (TEC) maps demonstrate lower reconstruction errors than competing methods.

By Jingyang Li, Michael K. Ng
arXiv Machine Learning
4d ago

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.

By Victor L\'eger, Florent Chatelain
arXiv Machine Learning
Jun 16

Unsupervised Learning for Missing Modalities in Multimodal Learning

arXiv:2606. 15743v1 Announce Type: new Abstract: This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature embeddings in a task-independent manner before supervised prediction.

By Hassan Ismkhan, Hamid Bouchahcia