The paper systematically evaluates how five dimensionality reduction methods—PCA, Kernel PCA, VAE, Isomap, and MDS—affect the performance of four clustering algorithms (k‑means, AHC, GMM, and OPTICS). Using the Adjusted Rand Index, the study compares clustering quality with and without dimensionality reduction at levels of k‑1, 25%, and 50% of the original dimensions. Results highlight that the choice of reduction technique and its level must be carefully matched to the data’s geometry and the clustering algorithm used.
By Ousmane Assani Amate, Elyes Lounissi, Mohammadreza Bakhtyari, \'Emilie Roy, Roman Sarrazin-Gendron, Vladimir Makarenkov
arXiv:2505.18918v4 Announce Type: replace-cross
Abstract: Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend...
By Javier Salazar Cavazos, Jeffrey A Fessler, Laura Balzano
arXiv:2608. 15313v1 Announce Type: cross Abstract: In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and dimensionality reduction that incorporates differential geometric information into the covariance structure of classical PCA.
By Alexandre L. M. Levada
arXiv:2607. 09490v1 Announce Type: cross Abstract: Terminal embeddings have emerged as a powerful tool for dimension reduction.
By Alexander Munteanu, Matteo Russo, David Saulpic, Chris Schwiegelshohn
arXiv:2609.00647v1 Announce Type: new
Abstract: Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. Its pointwise optimization, howeve...
By Xiaoyu Lian, Yuchao Zhang, Shuyin Xia, Siqi Zhong, Xuzhao Xiang
arXiv:2609.06468v1 Announce Type: new
Abstract: K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its...
By Abhiyan Dhakal (Kathmandu University), Pranish Kafle (Kathmandu University), Rajani Chulyadyo (Kathmandu University)
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2608.30093v1 Announce Type: cross
Abstract: We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) me...
By Anirban Mondal, Paromita Banerjee, Abhijit Mandal
arXiv:2607. 24405v1 Announce Type: new Abstract: In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data.
By Abdallah Alabdallah
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
By Savik Kinger, Luciano Dyballa, Steven W. Zucker
arXiv:2505. 04346v2 Announce Type: replace Abstract: Clustering aims at partitioning data points into groups of similar objects without knowing about the class labels.
By Arghya Pratihar, Kushal Bose, Swagatam Das