arXiv Machine Learning

Federated Clustering with Unknown Local and Global Cluster Cardinalities

arXiv Statistics ML
6d ago

Ensembles of Exactly Solved Subsamples for Clusterwise Regression: Trimming Without a Trimming Level

The paper proposes an ensemble method for clusterwise regression that uses exact solutions on many small random subsamples. Each subsample is solved to global optimality, extended to the full data via nearest-surface assignment, and the resulting partitions are combined by voting or selection. The method achieves high accuracy even with up to 20% gross outliers and can estimate the trimming level without prior knowledge, outperforming traditional trimmed alternation in worst‑case scenarios.

By Samir Orujov
arXiv Machine Learning
Jun 19

Variational Consensus Monte Carlo for Bayesian Mixture

arXiv:2606. 19643v1 Announce Type: cross Abstract: Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.

By Julie Fendler, Francesca L. Crowe, Tom Marshall, Sylvia Richardson, Paul D. W. Kirk
arXiv Machine Learning
Sep 17

Bracketing Uncertainty in Clustering Under the Manifold Hypothesis

The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.

By Savik Kinger, Luciano Dyballa, Steven W. Zucker
arXiv Machine Learning
1d ago

StoCFL: A Stochastically Clustered Federated Learning Framework for Non-IID Data with Dynamic Client Participation

StoCFL is a clustered federated learning framework designed to address Non-IID data and dynamic client participation. It introduces a flexible clustering mechanism that allows arbitrary client participation and accommodates newly joined clients, improving data efficiency and model performance. Experiments on four Non-IID settings and a real-world dataset demonstrate that StoCFL achieves promising cluster results even when the number of clusters is unknown, outperforming baseline approaches across various scenarios.

By Dun Zeng, Xiangjing Hu, Shiyu Liu, Yue Yu, Qifan Wang, Zenglin Xu
arXiv Machine Learning
Sep 18

Federated Soft Clustering via Generalized Total Variation Minimization

The paper introduces federated soft clustering for devices in a federated learning network, each fitting a personalized Gaussian mixture model. It proposes Generalized Total Variation Minimization (GTVMin) to couple local maximum likelihood problems via a graph regularizer that penalizes discrepancies between connected nodes’ models. Three discrepancy measures are compared: a squared Euclidean distance requiring component matching, a Monte‑Carlo approximated Kullback‑Leibler divergence, and a closed‑form maximum mean discrepancy; all are optimized with synchronous projected gradient updates, with a convergence guarantee for the smooth MMD instance.

By Shamsiiat Abdurakhmanova, Alexander Jung