arXiv:2606. 29261v1 Announce Type: new Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separation (BSS) to represent a D-dimensional observation vector as an unknown multivariate nonlinear mapping of C latent variables.
By Ivica Kopriva
The paper systematically evaluates how five dimensionality reduction methods—PCA, Kernel PCA, VAE, Isomap, and MDS—affect the performance of four clustering algorithms (k‑means, AHC, GMM, and OPTICS). Using the Adjusted Rand Index, the study compares clustering quality with and without dimensionality reduction at levels of k‑1, 25%, and 50% of the original dimensions. Results highlight that the choice of reduction technique and its level must be carefully matched to the data’s geometry and the clustering algorithm used.
By Ousmane Assani Amate, Elyes Lounissi, Mohammadreza Bakhtyari, \'Emilie Roy, Roman Sarrazin-Gendron, Vladimir Makarenkov
arXiv:2606. 08322v1 Announce Type: new Abstract: To characterize the US airline profit cycles from 1995 to 2020, the authors of Renold et al.
By Andreas Schlapbach
SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.
By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2608. 14215v1 Announce Type: new Abstract: Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains.
By Johanna Hillebrand, Jan H\"ockendorff, J\"urgen Kusche, Kelin Luo, Heiko R\"oglin, Melanie Schmidt, Christian Sohler, Bernd Uebbing
arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.
By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters
arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.
By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw
arXiv:2608. 15313v1 Announce Type: cross Abstract: In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and dimensionality reduction that incorporates differential geometric information into the covariance structure of classical PCA.
By Alexandre L. M. Levada
The paper introduces two randomized approaches to accelerate spectral co‑clustering of word‑document matrices: one based on randomized SVD via random projection, and another combining partial SVD with element‑wise random sampling. Experiments on real and synthetic data show both methods cut runtime compared to full SVD, with the projection technique offering more consistent performance across varied sparsity levels, while the sampling method excels on denser matrices. The study highlights that the choice of approximation should align with the data’s structural properties.
By Fateme Mazdarani, Carlos Toxtli
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
By Filip Kosiorowski, Grzegorz Sroka
arXiv:2212. 07944v4 Announce Type: replace Abstract: We study a multi-factor block model for variable clustering and connect it to regularized subspace clustering through a distributionally robust version of nodewise regression.
By Kaizheng Wang, Xiao Xu, Xun Yu Zhou