arXiv Machine Learning
Jun 30

Nonlinear mixture model motivated subspace clustering

arXiv:2606. 29261v1 Announce Type: new Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separation (BSS) to represent a D-dimensional observation vector as an unknown multivariate nonlinear mapping of C latent variables.

By Ivica Kopriva
arXiv Machine Learning
Sep 24

Assessing the impact of dimensionality reduction on clustering performance - a systematic study

The paper systematically evaluates how five dimensionality reduction methods—PCA, Kernel PCA, VAE, Isomap, and MDS—affect the performance of four clustering algorithms (k‑means, AHC, GMM, and OPTICS). Using the Adjusted Rand Index, the study compares clustering quality with and without dimensionality reduction at levels of k‑1, 25%, and 50% of the original dimensions. Results highlight that the choice of reduction technique and its level must be carefully matched to the data’s geometry and the clustering algorithm used.

By Ousmane Assani Amate, Elyes Lounissi, Mohammadreza Bakhtyari, \'Emilie Roy, Roman Sarrazin-Gendron, Vladimir Makarenkov
arXiv Machine Learning
Sep 23

SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA

SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.

By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park