arXiv:2607. 01945v1 Announce Type: cross Abstract: The classical $k$-means clustering cannot be directly used to incomplete data, and existing $k$-means-based clustering for missing data primarily focus on improving the practical accuracy of clustering, whereas most of them lack theoretical guarantees in the asymptotic sense.
By Xin Guan
arXiv:2609.00616v1 Announce Type: cross
Abstract: Matrix-variate data with missing entries arise frequently in applications where observations are naturally organized as two-dimensional arrays. Altho...
By Hanzhang Lu, Jeffrey L. Andrews, Ryan P. Browne
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
By Naitik Gada (Rochester Institute of Technology)
The paper introduces the Robust Graph Clustering Network for Multiple Missing Data (RGCN), a method designed to cluster graphs with simultaneous missing node attributes and structural links. RGCN employs a view‑decoupled dual‑branch imputation to reduce cross‑view interference, a multi‑hyperspherical mixture prior to improve cluster compactness and separability on a directional latent manifold, and a boundary‑aware contrastive enhancement objective to counteract cluster blurring caused by imputation bias. Experiments on real‑world datasets show that RGCN consistently outperforms state‑of‑the‑art baselines across various missing data patterns.
By Keyuan Qiu, Renda Han, Zhen Tang, Qiang He, Xingwei Wang, Wenxin Zhang, Guangzhen Yao, Junxin Chen, Qingjian Ni
arXiv:2411. 12438v2 Announce Type: replace-cross Abstract: We develop a new approach for clustering non-spherical (i.
By Prashanti Anderson, Mitali Bafna, Rares-Darius Buhai, Pravesh K. Kothari, David Steurer
arXiv:2508. 12450v2 Announce Type: replace Abstract: This article presents an adaptive mean shift algorithm in which every parameter used at a point is derived from that point's own distance distribution.
By \'Etienne Pepin
arXiv:2609.06468v1 Announce Type: new
Abstract: K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its...
By Abhiyan Dhakal (Kathmandu University), Pranish Kafle (Kathmandu University), Rajani Chulyadyo (Kathmandu University)
arXiv:2609.39613v1 Announce Type: new
Abstract: Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstr...
By Jinwei Li, Michelle Bruch, Daniel Tenbrinck
arXiv:2607. 01993v1 Announce Type: cross Abstract: The silhouette is one of the most widely used measures to assess the quality of a $k$-clustering of a dataset of $n$ elements.
By Ilie Sarpe, Federico Altieri, Andrea Pietracaprina, Geppino Pucci, Fabio Vandin
arXiv:2607. 24405v1 Announce Type: new Abstract: In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data.
By Abdallah Alabdallah
arXiv:2607. 06930v1 Announce Type: cross Abstract: Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis.
By Chuyao Zhang, E Li, Taochen Chen, Yiqun Zhang, Yuzhu Ji, Shuping Zhao, Peng Liu, Yiu-ming Cheung
arXiv:2607. 24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge.
By Filip Kosiorowski, Grzegorz Sroka