arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2502. 08397v3 Announce Type: replace-cross Abstract: Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together.
By Anna Livia Croella, Veronica Piccialli, Antonio M. Sudoso
arXiv:2607. 15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids.
By Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan
arXiv:2606. 31253v1 Announce Type: cross Abstract: The classical $k$-means clustering, based on distances computed from all data features, cannot be directly applied to incomplete data with missing values.
By Xin Guan
arXiv:2608.30093v1 Announce Type: cross
Abstract: We introduce a robust clustering method, MK-means DPD, that estimates cluster centers and covariance matrices using density power divergence (DPD) me...
By Anirban Mondal, Paromita Banerjee, Abhijit Mandal
arXiv:2607. 26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$.
By Vaclav Rozhon
arXiv:2607. 01945v1 Announce Type: cross Abstract: The classical $k$-means clustering cannot be directly used to incomplete data, and existing $k$-means-based clustering for missing data primarily focus on improving the practical accuracy of clustering, whereas most of them lack theoretical guarantees in the asymptotic sense.
By Xin Guan
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
By Dani\"el Bot, Leland McInnes, Jan Aerts
arXiv:2606. 10673v1 Announce Type: cross Abstract: Although some very common test beds exist for assessing the performance of clustering methods, large scale benchmarking is typically limited to relatively simplistic simulation set-ups.
By David P. Hofmeyr
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
By Filip Kosiorowski, Grzegorz Sroka
The paper "Mixed Data Clustering Survey and Challenges" discusses how the rise of big data has made clustering of heterogeneous datasets—containing both numerical and categorical variables—particularly difficult for traditional methods. It highlights the importance of hierarchical and explainable algorithms for producing interpretable results that aid decision‑making. The authors propose a new clustering approach based on pretopological spaces and benchmark it against classical numerical clustering algorithms and existing pretopological methods to evaluate its performance in the big data context.
By Maxence Choufa, Clement Cornet, Guillaume Guerard, Sonia Djebali, Loup-No\'e Levy