The paper introduces CluProp, a framework that treats varied‑density clustering in high‑dimensional spaces as a label propagation process over neighborhood graphs. By combining density‑based ideas with graph connectivity, it offers a deterministic propagation strategy that reduces parameter sensitivity and scales efficiently to millions of points. CluProp is agnostic to distance metrics and consistently outperforms existing baselines in accuracy while processing large datasets in minutes.
By Yingtao Zheng, Hugo Phibbs, Ninh Pham
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
arXiv:2605. 22410v2 Announce Type: replace Abstract: Spectral clustering largely depends on the affinity graph, yet constructing a graph that preserves reliable local connectivity while adapting to heterogeneous data structures remains challenging.
By Zeqiang Xian, Caihui Liu, Yong Zhang, Wenjing Qiu
arXiv:2512. 16558v3 Announce Type: replace Abstract: Clustering is a cornerstone of modern data analysis.
By Dani\"el Bot, Leland McInnes, Jan Aerts
The paper introduces Inductive Correlation Clustering, a new framework that uses Graph Neural Networks to solve the Correlation Clustering problem on unseen graph instances. By learning common structural patterns and node features, the method generalizes to new graphs with minimal computational overhead, achieving inference times up to five orders of magnitude faster while maintaining an approximation ratio within about 10% of the best baseline. It also demonstrates competitive performance on standard transductive benchmarks and serves as an efficient learnable pooling layer for graph classification tasks.
By Francesco Paolo Nerini, Francesco Bonchi, Arijit Khan, Andr\'e Panisson
arXiv:2508. 02989v2 Announce Type: replace Abstract: We propose a novel perspective on varied-density clustering for high-dimensional data by framing it as a label propagation process in neighborhood graphs that adapt to local density variations.
By Ninh Pham, Yingtao Zheng, Hugo Phibbs
The paper introduces the Universal Clustering Problem (UCP), a framework that captures the optimisation core common to many clustering methods by maximizing a polynomial‑time computable partition utility over a finite metric space. It proves UCP is NP‑hard through reductions from graph colouring and exact cover by 3‑sets, showing that popular algorithms such as k‑means, GMMs, DBSCAN, spectral clustering, and affinity propagation inherit this intractability. The authors argue that this unified hardness explains typical failure modes—like local optima and greedy merge traps—and suggest moving toward stability‑aware objectives and interaction‑driven formulations with explicit guarantees.
By Angshul Majumdar
arXiv:2609.26063v1 Announce Type: new
Abstract: Federated graph learning (FGL) enables multiple clients to collaboratively train graph models without sharing their private graph data, providing a pro...
By Yinlin Zhu, Di Wu, Wang Luo, Guocong Quan, Miao Hu
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2607. 05464v1 Announce Type: cross Abstract: The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects.
By Yiqun Zhang, Yiu-ming Cheung
arXiv:2609.26748v1 Announce Type: cross
Abstract: Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters,...
By Siyi Wang, Alexandre Leblanc, Paul D. McNicholas
arXiv:2607. 24237v1 Announce Type: new Abstract: Many existing clustering methods are designed based on a set-oriented definition---a cluster is a set of similar points---relying a point-to-point similarity function to find similar points.
By Kai Ming Ting, Kaifeng Zhang, Sanjay Chawla