arXiv Machine Learning

A novel k-means clustering approach using two distance measures for Gaussian data

arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.

arXiv Machine Learning
Aug 28

Absolute indices for determining compactness, separability and number of clusters

The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.

By Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri
arXiv Machine Learning
Sep 24

Assessing the impact of dimensionality reduction on clustering performance - a systematic study

The paper systematically evaluates how five dimensionality reduction methods—PCA, Kernel PCA, VAE, Isomap, and MDS—affect the performance of four clustering algorithms (k‑means, AHC, GMM, and OPTICS). Using the Adjusted Rand Index, the study compares clustering quality with and without dimensionality reduction at levels of k‑1, 25%, and 50% of the original dimensions. Results highlight that the choice of reduction technique and its level must be carefully matched to the data’s geometry and the clustering algorithm used.

By Ousmane Assani Amate, Elyes Lounissi, Mohammadreza Bakhtyari, \'Emilie Roy, Roman Sarrazin-Gendron, Vladimir Makarenkov