arXiv Machine Learning

Sector-Mean: Deterministic Initialization of K-Means Centroids via Angular Sector Partitioning

arXiv Machine Learning
Jul 20

Data-Native Global Optimization for Big Data K-means Clustering

arXiv:2607. 15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids.

By Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan
arXiv Machine Learning
Jul 30

Randomizing the Number of Centers in k-means++

arXiv:2607. 26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$.

By Vaclav Rozhon
arXiv AI
Sep 4

Mixed Data Clustering Survey and Challenges

The paper "Mixed Data Clustering Survey and Challenges" discusses how the rise of big data has made clustering of heterogeneous datasets—containing both numerical and categorical variables—particularly difficult for traditional methods. It highlights the importance of hierarchical and explainable algorithms for producing interpretable results that aid decision‑making. The authors propose a new clustering approach based on pretopological spaces and benchmark it against classical numerical clustering algorithms and existing pretopological methods to evaluate its performance in the big data context.

By Maxence Choufa, Clement Cornet, Guillaume Guerard, Sonia Djebali, Loup-No\'e Levy