arXiv Machine Learning

Scalable Varied-Density Clustering via Graph Propagation

arXiv:2508. 02989v2 Announce Type: replace Abstract: We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations.

arXiv Machine Learning
Aug 27

Towards Robust and Scalable Density-based Clustering via Graph Propagation

The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.

By Yingtao Zheng, Hugo Phibbs, Ninh Pham
arXiv Statistics ML
2d ago

Gradient-Guided Density Peak Clustering

Gradient-Guided Density Peak Clustering (GGDPC) enhances traditional density peak clustering by performing a gradient ascent step before each nearest‑neighbor uphill search, aiming to stabilize uphill paths in low‑density regions. The authors develop a stability theory linking the GGDPC graph to the gradient ascent flow of the population density, and establish consistency across five criteria: recovery of local modes, adjusted Rand index, dendrogram (cluster tree), path length, and waterfall measure. These results offer new statistical, geometric, and topological insights into DPC‑type clustering algorithms.

By Yikun Zhang, Yen-Chi Chen
arXiv Machine Learning
Sep 3

From topology learning to graph generation: A unifying perspective

The article reviews the problem of learning graph structures from data, noting that research has traditionally split into two paths: inferring the topology of a single graph from observations on it, and learning a generative distribution from multiple observed graphs to sample new ones. It proposes a unified framework that treats both as inverse problems of a common graph generation process, reviews key methods, and discusses their interrelations, strengths, and limitations. The review highlights opportunities for cross‑paradigm integration and outlines future research directions.

By Xiaowen Dong, Hoi-To Wai, Siheng Chen, Laura Toni, Dorina Thanou
arXiv Machine Learning
Sep 11

SynCo: Synthetic Community-Aware Attributed Graph Generator for Graph Neural Network Benchmarking

SynCo is a synthetic graph generator that lets users control node degree distributions and sub‑community structures, addressing limitations of existing generators that rely on power‑law distributions and lack flexibility. It is evaluated on graph mimicking, hyperparameter tuning, and node clustering, outperforming state‑of‑the‑art methods while preserving original data distributions. SynCo can generate large graphs with up to 2.1 million nodes.

By Guilherme Henrique Messias, Mariana Caravanti de Souza, Sylvia Iasulaitis, Alan Dem\'etrius Baria Valejo
arXiv Machine Learning
Aug 20

Enhancing Distance-Based Graph Autoencoders with Structural Penalties for Dynamic Graph Embedding

The paper introduces three distance‑based graph autoencoder variants that add structural penalties to the reconstruction loss. All models use a two‑layer Graph Convolutional Network encoder and a Euclidean‑distance decoder, with two node‑level regularizers: a hub penalty based on degree centrality and a penalty based on Natural Community Local Intrinsic Dimensionality (NC‑LID). Experiments on multiple dynamic graph datasets show that incorporating NC‑LID regularization consistently improves reconstruction performance compared to baselines without structural regularization and to the hub‑aware variant.

By Aleksandar Tom\v{c}i\'c, Milo\v{s} Savi\'c, Milo\v{s} Radovanovi\'c