Consider the following variation on the Hierarchical Clustering problem: Usually, while building a hierarchical clustering, one recursively partitions the data until each cluster becomes a singleton. We relax the halting condition of the recursive process to stop whenever the remaining cluster is a graph belonging to a class $\mathcal{F}$.
arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2504. 19419v3 Announce Type: replace Abstract: Local clustering aims to identify specific substructures within a large graph without any additional structural information of the graph.
By Zhaiming Shen, Sung Ha Kang
arXiv:2607. 13217v1 Announce Type: cross Abstract: Consider the following variation on the Hierarchical Clustering problem: Usually, while building a hierarchical clustering, one recursively partitions the data until each cluster becomes a singleton.
By Micha{\l} Szyfelbein, Dariusz Dereniowski
arXiv:2602. 08542v3 Announce Type: replace-cross Abstract: Given a weighted undirected graph, a number of clusters $k$, and an exponent $z$, the goal in the $(k, z)$-clustering problem on graphs is to select $k$ vertices as centers that minimize the sum of the distances raised to the power $z$ of each vertex to its closest center.
By Emilio Cruciani, Sebastian Forster, Antonis Skarlatos
arXiv:2508. 02158v2 Announce Type: replace-cross Abstract: Detection of planted subgraphs in Erd\"os-R\'enyi random graphs has been extensively studied, leading to a rich body of results characterizing both statistical and computational thresholds.
By Dor Elimelech, Wasim Huleihel
arXiv:2412. 03008v2 Announce Type: replace-cross Abstract: Local/seeded clustering aims to find a compact cluster near the given starting instances.
By Zihao Li, Dongqi Fu, Hengyu Liu, Jingrui He
arXiv:2606. 14335v1 Announce Type: cross Abstract: Recovering structural information from noisy high-dimensional data is a fundamental task in statistical inference.
By Zhe Hou, Jingcheng Liu
The paper introduces the Universal Clustering Problem (UCP), a framework that captures the optimisation core common to many clustering methods by maximizing a polynomial‑time computable partition utility over a finite metric space. It proves UCP is NP‑hard through reductions from graph colouring and exact cover by 3‑sets, showing that popular algorithms such as k‑means, GMMs, DBSCAN, spectral clustering, and affinity propagation inherit this intractability. The authors argue that this unified hardness explains typical failure modes—like local optima and greedy merge traps—and suggest moving toward stability‑aware objectives and interaction‑driven formulations with explicit guarantees.
By Angshul Majumdar
The paper presents provable guarantees for a spectral method that recovers binary node labels on signed graphs with edge‑flip noise. It provides graph‑structure‑agnostic bounds on approximate inference accuracy and maximum angle deviation, using matrix concentration and eigenvector perturbation techniques. The results connect to the Cheeger constant and are validated with synthetic experiments, marking the first theoretical analysis of this spectral approach.
By Violet Zheng, Jean Honorio
The paper introduces a data‑driven method for learning Random Geometric Graphs (RGGs) in probabilistic metric spaces. It defines a distance function based on the cumulative distribution of a disparity variable that captures differences in vertex connectivity and correlation of attached random variables, enabling edges to exist with a specified probability. The approach includes a rejection‑sampling technique for edge probability estimation and a closed‑form posterior for learning the inter‑observable correlation matrix, and it is demonstrated on highly multivariate real datasets.
By Dalia Chakrabarty, Kangrui Wang, Chuqiao Zhang, Ye Liu
arXiv:2510. 15076v2 Announce Type: replace Abstract: The $\ell_p$-norm objectives for correlation clustering present a fundamental trade-off between minimizing total disagreements (the $\ell_1$-norm) and ensuring fairness to individual nodes (the $\ell_\infty$-norm).
By Sami Davies, Benjamin Moseley, Heather Newman