arXiv Machine Learning

Variable Clustering via Distributionally Robust Nodewise Regression

arXiv:2212. 07944v4 Announce Type: replace Abstract: We study a multi-factor block model for variable clustering and connect it to regularized subspace clustering through a distributionally robust version of nodewise regression.

arXiv Machine Learning
Jun 30

Nonlinear mixture model motivated subspace clustering

arXiv:2606. 29261v1 Announce Type: new Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separation (BSS) to represent a D-dimensional observation vector as an unknown multivariate nonlinear mapping of C latent variables.

By Ivica Kopriva
arXiv Statistics ML
Aug 31

Robust model-based clustering via mixtures of multivariate pseudo-Voigt distributions

The paper introduces a multivariate pseudo‑Voigt mixture model, combining Gaussian and Cauchy components with shared location and scale parameters, for robust clustering and outlier detection. Parameter estimation is performed using an EM algorithm that leverages latent variables for efficient likelihood inference. The authors evaluate the model through simulations and real data, comparing it to established robust mixtures such as contaminated normals, and demonstrate its effectiveness on heavy‑tailed datasets.

By Babak F. Dehkordi, Jeffrey L. Andrews, Andrew Jirasek
arXiv Machine Learning
Sep 18

Federated Soft Clustering via Generalized Total Variation Minimization

The paper introduces federated soft clustering for devices in a federated learning network, each fitting a personalized Gaussian mixture model. It proposes Generalized Total Variation Minimization (GTVMin) to couple local maximum likelihood problems via a graph regularizer that penalizes discrepancies between connected nodes’ models. Three discrepancy measures are compared: a squared Euclidean distance requiring component matching, a Monte‑Carlo approximated Kullback‑Leibler divergence, and a closed‑form maximum mean discrepancy; all are optimized with synchronous projected gradient updates, with a convergence guarantee for the smooth MMD instance.

By Shamsiiat Abdurakhmanova, Alexander Jung
arXiv Machine Learning
Sep 3

Clustering Three-Way Data with Outliers

The paper introduces a method for clustering matrix-variate normal data that accounts for outliers. It extends the OCLUST algorithm by employing subset log-likelihood distributions and an iterative trimming procedure. This approach enables robust clustering of complex structured data such as images and time series.

By Katharine M. Clark, Paul D. McNicholas