arXiv Statistics ML
2d ago

Gradient-Guided Density Peak Clustering

Gradient-Guided Density Peak Clustering (GGDPC) enhances traditional density peak clustering by performing a gradient ascent step before each nearest‑neighbor uphill search, aiming to stabilize uphill paths in low‑density regions. The authors develop a stability theory linking the GGDPC graph to the gradient ascent flow of the population density, and establish consistency across five criteria: recovery of local modes, adjusted Rand index, dendrogram (cluster tree), path length, and waterfall measure. These results offer new statistical, geometric, and topological insights into DPC‑type clustering algorithms.

By Yikun Zhang, Yen-Chi Chen
arXiv Machine Learning
2d ago

T-ARC: Topology-Aware Randomized Clustering via Distributionally Robust Stochastic Block Models

The paper introduces T-ARC, a clustering algorithm that integrates topological information into the K‑means objective by coupling a data‑fidelity term with a graph‑cut penalty. The latent graph is modeled as a random realization from a Stochastic Block Model, whose parameter is optimized via Distributionally Robust Optimization, using a persistence‑based similarity matrix derived from zero‑dimensional persistent homology. Experiments on synthetic non‑convex data and Fashion‑MNIST subsets demonstrate that T‑ARC recovers latent topological structures and outperforms K‑means on curved and interleaved clusters while remaining competitive and more stable on real data.

By Serena Grazia De Benedictis, Andersen Ang, Nicoletta Del Buono, Flavia Esposito, Laura Selicato
arXiv Machine Learning
Jul 30

Randomizing the Number of Centers in k-means++

arXiv:2607. 26202v1 Announce Type: cross Abstract: The $k$-means++ algorithm is a standard and widely used seeding method for $k$-means clustering, but for a fixed number $k$ of centers its worst-case expected approximation ratio is $\Theta(\log k)$.

By Vaclav Rozhon