The K-SCAN Clustering Algorithm
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
arXiv:2505. 04346v2 Announce Type: replace Abstract: Clustering aims at partitioning data points into groups of similar objects without knowing about the class labels.
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
arXiv:2609.26748v1 Announce Type: cross Abstract: Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters,...
The paper introduces absolute cluster indices that assess both compactness and separability of clusters, moving beyond relative measures commonly used in clustering validation. It defines a compactness function for each cluster and a set of neighboring points for cluster pairs to evaluate cluster quality and overall distribution margin. These indices are applied to determine the true number of clusters and are compared against widely-used validity indices on synthetic and real-world datasets.
Gradient-Guided Density Peak Clustering (GGDPC) enhances traditional density peak clustering by performing a gradient ascent step before each nearest‑neighbor uphill search, aiming to stabilize uphill paths in low‑density regions. The authors develop a stability theory linking the GGDPC graph to the gradient ascent flow of the population density, and establish consistency across five criteria: recovery of local modes, adjusted Rand index, dendrogram (cluster tree), path length, and waterfall measure. These results offer new statistical, geometric, and topological insights into DPC‑type clustering algorithms.
arXiv:2606. 00327v1 Announce Type: cross Abstract: Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries.
The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.
arXiv:2604. 18801v2 Announce Type: replace Abstract: Scientific particle simulations in cosmology, molecular dynamics, and fluid dynamics produce large-scale datasets whose storage, movement, and analysis increasingly rely on lossy compression.
arXiv:2605. 30225v2 Announce Type: replace Abstract: Clustering is an unsupervised technique for grouping data points by similarity.
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.