arXiv Machine Learning

An unsupervised clustering analysis of breast cancer data derived from electronic health records enhanced through UMAP dimensionality reduction

arXiv:2607. 19089v1 Announce Type: new Abstract: Breast cancer is one of the most widespread types of cancer, affecting approximately 8 million women worldwide.

arXiv Statistics ML
Sep 23

Efficient and scalable clustering of survival curves

arXiv:2512.16481v2 Announce Type: replace-cross Abstract: Survival analysis encompasses a broad range of methods for analyzing time-to-event data, with one key objective being the comparison of survi...

By Nora M. Villanueva, Marta Sestelo, Luis Meira-Machado
arXiv Machine Learning
Sep 24

Assessing the impact of dimensionality reduction on clustering performance - a systematic study

The paper systematically evaluates how five dimensionality reduction methods—PCA, Kernel PCA, VAE, Isomap, and MDS—affect the performance of four clustering algorithms (k‑means, AHC, GMM, and OPTICS). Using the Adjusted Rand Index, the study compares clustering quality with and without dimensionality reduction at levels of k‑1, 25%, and 50% of the original dimensions. Results highlight that the choice of reduction technique and its level must be carefully matched to the data’s geometry and the clustering algorithm used.

By Ousmane Assani Amate, Elyes Lounissi, Mohammadreza Bakhtyari, \'Emilie Roy, Roman Sarrazin-Gendron, Vladimir Makarenkov
arXiv AI
Jun 30

Interpretable Clustering: A Survey

arXiv:2409. 00743v4 Announce Type: replace-cross Abstract: In recent years, much of the research on clustering algorithms has primarily focused on enhancing their accuracy and efficiency, frequently at the expense of interpretability.

By Lianyu Hu, Mudi Jiang, Junjie Dong, Xinying Liu, Zengyou He
arXiv Machine Learning
1d ago

Classification Based on Association Rules Algorithm for Breast Cancer

The paper presents a new association rule-based data mining technique for classifying breast cancer, emphasizing early detection. It introduces a weighted classification approach that uses three core algorithms: Rule Generation, Rule Pruning, and Rule Prediction. The method identifies frequent itemsets, prunes rules into major and minor groups, and applies the pruned rules to classify test data, demonstrating feasibility and performance on several samples.

By Ali Alsalama, Ahmed Kubba, Ghaith Jamjoum, Zaher Al Aghbari
arXiv Machine Learning
Aug 28

Absolute indices for determining compactness, separability and number of clusters

The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.

By Adil M. Bagirov, Ramiz M. Aliguliyev, Nargiz Sultanova, Sona Taheri