arXiv Machine Learning

A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment

arXiv:2608. 14968v1 Announce Type: cross Abstract: We consider nonparametric regression when the association between a response and its covariates changes across an unknown partition of a spatial domain.

arXiv AI
Sep 21

Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data

The paper introduces FedDCN, a federated deep clustering network that jointly optimizes reconstruction and clustering losses for high‑dimensional, heterogeneous data. It addresses challenges of non‑IID client data by generating synthetic augmentations and applying geometric regularization to align latent spaces. Experiments show the method’s effectiveness under both IID and non‑IID settings, and the authors outline future research directions.

By Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik, Anna Wilbik
arXiv Machine Learning
Aug 26

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.

By Kai Ming Ting, Wei-Jie Xu, Hang Zhang
arXiv Machine Learning
5d ago

DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology

DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.

By Joshua Dimasaka, Christian Gei{\ss}, Emily So
arXiv Machine Learning
Sep 25

Selective Inference for Deep Clustering in Latent Spaces

The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.

By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
arXiv Machine Learning
Sep 18

Federated Soft Clustering via Generalized Total Variation Minimization

The paper introduces federated soft clustering for devices in a federated learning network, each fitting a personalized Gaussian mixture model. It proposes Generalized Total Variation Minimization (GTVMin) to couple local maximum likelihood problems via a graph regularizer that penalizes discrepancies between connected nodes’ models. Three discrepancy measures are compared: a squared Euclidean distance requiring component matching, a Monte‑Carlo approximated Kullback‑Leibler divergence, and a closed‑form maximum mean discrepancy; all are optimized with synchronous projected gradient updates, with a convergence guarantee for the smooth MMD instance.

By Shamsiiat Abdurakhmanova, Alexander Jung