arXiv Machine Learning

Towards Robust and Scalable Density-based Clustering via Graph Propagation

The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.

arXiv Machine Learning
Jul 13

Scalable Varied-Density Clustering via Graph Propagation

arXiv:2508. 02989v2 Announce Type: replace Abstract: We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations.

By Ninh Pham, Yingtao Zheng, Hugo Phibbs
arXiv Statistics ML
2d ago

Gradient-Guided Density Peak Clustering

Gradient-Guided Density Peak Clustering (GGDPC) enhances traditional density peak clustering by performing a gradient ascent step before each nearest‑neighbor uphill search, aiming to stabilize uphill paths in low‑density regions. The authors develop a stability theory linking the GGDPC graph to the gradient ascent flow of the population density, and establish consistency across five criteria: recovery of local modes, adjusted Rand index, dendrogram (cluster tree), path length, and waterfall measure. These results offer new statistical, geometric, and topological insights into DPC‑type clustering algorithms.

By Yikun Zhang, Yen-Chi Chen
arXiv Machine Learning
Sep 3

From topology learning to graph generation: A unifying perspective

The article reviews the problem of learning graph structures from data, noting that research has traditionally split into two paths: inferring the topology of a single graph from observations on it, and learning a generative distribution from multiple observed graphs to sample new ones. It proposes a unified framework that treats both as inverse problems of a common graph generation process, reviews key methods, and discusses their interrelations, strengths, and limitations. The review highlights opportunities for cross‑paradigm integration and outlines future research directions.

By Xiaowen Dong, Hoi-To Wai, Siheng Chen, Laura Toni, Dorina Thanou
arXiv Machine Learning
Aug 31

Optimal Transport for Network Comparison: A Review with Machine Learning Applications

The paper reviews the use of optimal transport for comparing undirected, unweighted graphs, focusing on three main distances: Wasserstein, Gromov-Wasserstein, and Bures-Wasserstein. It discusses closed-form solutions for the Wasserstein distance in one dimension, how transport plans identify influential nodes after perturbations, and derives spectral bounds for the Bures-Wasserstein distance to avoid full decompositions. The authors evaluate these distances on synthetic clustering data and a real-world time‑series network for anomaly detection.

By James Hyun, Fran\c{c}ois G. Meyer