The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.
By Samir Orujov
arXiv:2607. 01945v1 Announce Type: cross Abstract: The classical $k$-means clustering cannot be directly used to incomplete data, and existing $k$-means-based clustering for missing data primarily focus on improving the practical accuracy of clustering, whereas most of them lack theoretical guarantees in the asymptotic sense.
By Xin Guan
arXiv:2606. 19643v1 Announce Type: cross Abstract: Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.
By Julie Fendler, Francesca L. Crowe, Tom Marshall, Sylvia Richardson, Paul D. W. Kirk
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
By Savik Kinger, Luciano Dyballa, Steven W. Zucker
arXiv:2411. 01576v3 Announce Type: replace Abstract: The explainable clustering problem was first posed by Moshkovitz et al.
By Maximilian Fleissner, Maedeh Zarvandi, Debarghya Ghoshdastidar
arXiv:2506. 22427v2 Announce Type: replace-cross Abstract: We propose CLoVE (Clustering of Loss Vector Embeddings), a novel algorithm for Clustered Federated Learning (CFL).
By Randeep Bhatia, Nikos Papadis, Murali Kodialam, TV Lakshman, Sayak Chakrabarty
StoCFL is a clustered federated learning framework designed to address Non-IID data and dynamic client participation. It introduces a flexible clustering mechanism that allows arbitrary client participation and accommodates newly joined clients, improving data efficiency and model performance. Experiments on four Non-IID settings and a real-world dataset demonstrate that StoCFL achieves promising cluster results even when the number of clusters is unknown, outperforming baseline approaches across various scenarios.
By Dun Zeng, Xiangjing Hu, Shiyu Liu, Yue Yu, Qifan Wang, Zenglin Xu
arXiv:2505. 20532v2 Announce Type: replace Abstract: This paper studies robust one-shot aggregation for distributed and federated Independent Component Analysis (ICA).
By Dian Jin, Xin Bing, Yuqian Zhang
arXiv:2609.06468v1 Announce Type: new
Abstract: K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its...
By Abhiyan Dhakal (Kathmandu University), Pranish Kafle (Kathmandu University), Rajani Chulyadyo (Kathmandu University)
arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2508. 12450v2 Announce Type: replace Abstract: This article presents an adaptive mean shift algorithm in which every parameter used at a point is derived from that point's own distance distribution.
By \'Etienne Pepin
The paper introduces federated soft clustering for devices in a federated learning network, each fitting a personalized Gaussian mixture model. It proposes Generalized Total Variation Minimization (GTVMin) to couple local maximum likelihood problems via a graph regularizer that penalizes discrepancies between connected nodes’ models. Three discrepancy measures are compared: a squared Euclidean distance requiring component matching, a Monte‑Carlo approximated Kullback‑Leibler divergence, and a closed‑form maximum mean discrepancy; all are optimized with synchronous projected gradient updates, with a convergence guarantee for the smooth MMD instance.
By Shamsiiat Abdurakhmanova, Alexander Jung