The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.
By Yingtao Zheng, Hugo Phibbs, Ninh Pham
Gradient-Guided Density Peak Clustering (GGDPC) enhances traditional density peak clustering by performing a gradient ascent step before each nearest‑neighbor uphill search, aiming to stabilize uphill paths in low‑density regions. The authors develop a stability theory linking the GGDPC graph to the gradient ascent flow of the population density, and establish consistency across five criteria: recovery of local modes, adjusted Rand index, dendrogram (cluster tree), path length, and waterfall measure. These results offer new statistical, geometric, and topological insights into DPC‑type clustering algorithms.
By Yikun Zhang, Yen-Chi Chen
arXiv:2608. 06990v1 Announce Type: cross Abstract: Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning.
By Yuning Yu, Jos\'e Rodr\'iguez-Pi\~neiro, Xuefeng Yin, Bin Feng
The article reviews the problem of learning graph structures from data, noting that research has traditionally split into two paths: inferring the topology of a single graph from observations on it, and learning a generative distribution from multiple observed graphs to sample new ones. It proposes a unified framework that treats both as inverse problems of a common graph generation process, reviews key methods, and discusses their interrelations, strengths, and limitations. The review highlights opportunities for cross‑paradigm integration and outlines future research directions.
By Xiaowen Dong, Hoi-To Wai, Siheng Chen, Laura Toni, Dorina Thanou
arXiv:2505. 21285v5 Announce Type: replace Abstract: This work proposes a framework LGKDE that learns kernel density estimation for graphs.
By Xudong Wang, Ziheng Sun, Chris Ding, Jicong Fan
arXiv:2504. 19419v3 Announce Type: replace Abstract: Local clustering aims to identify specific substructures within a large graph without any additional structural information of the graph.
By Zhaiming Shen, Sung Ha Kang
arXiv:2607. 05469v1 Announce Type: cross Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks.
By Jingyun Zhang, Hao Peng, Jianxin Li, Angsheng Li, Philip S. Yu
SynCo is a synthetic graph generator that lets users control node degree distributions and sub‑community structures, addressing limitations of existing generators that rely on power‑law distributions and lack flexibility. It is evaluated on graph mimicking, hyperparameter tuning, and node clustering, outperforming state‑of‑the‑art methods while preserving original data distributions. SynCo can generate large graphs with up to 2.1 million nodes.
By Guilherme Henrique Messias, Mariana Caravanti de Souza, Sylvia Iasulaitis, Alan Dem\'etrius Baria Valejo
The paper introduces three distance‑based graph autoencoder variants that add structural penalties to the reconstruction loss. All models use a two‑layer Graph Convolutional Network encoder and a Euclidean‑distance decoder, with two node‑level regularizers: a hub penalty based on degree centrality and a penalty based on Natural Community Local Intrinsic Dimensionality (NC‑LID). Experiments on multiple dynamic graph datasets show that incorporating NC‑LID regularization consistently improves reconstruction performance compared to baselines without structural regularization and to the hub‑aware variant.
By Aleksandar Tom\v{c}i\'c, Milo\v{s} Savi\'c, Milo\v{s} Radovanovi\'c
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
By Dani\"el Bot, Leland McInnes, Jan Aerts
arXiv:2607. 08746v1 Announce Type: cross Abstract: While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally.
By Duen Horng Chau, Donghao Ren, Fred Hohman, Dominik Moritz
arXiv:2505. 04346v2 Announce Type: replace Abstract: Clustering aims at partitioning data points into groups of similar objects without knowing about the class labels.
By Arghya Pratihar, Kushal Bose, Swagatam Das