arXiv:2403.14830v2 Announce Type: replace
Abstract: Deep clustering partitions complex high-dimensional data using deep neural networks for clustering. It involves projecting data into lower-dimensio...
By Zeya Wang, Chenglong Ye
The paper introduces FedDCN, a federated deep clustering network that jointly optimizes reconstruction and clustering losses for high‑dimensional, heterogeneous data. It addresses challenges of non‑IID client data by generating synthetic augmentations and applying geometric regularization to align latent spaces. Experiments show the method’s effectiveness under both IID and non‑IID settings, and the authors outline future research directions.
By Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik, Anna Wilbik
arXiv:2609.26748v1 Announce Type: cross
Abstract: Clustering is an unsupervised learning technique that partitions unlabeled data into groups. Most existing methods require user-specified parameters,...
By Siyi Wang, Alexandre Leblanc, Paul D. McNicholas
arXiv:2609. 25605v1 Announce Type: cross Abstract: In this paper, we study the estimation of a marginal regression function from independent units with repeated binary, count, or continuous responses using ReLU deep neural networks.
By Kexuan Li
arXiv:2608. 00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy.
By Pulock Das, Yina Hou, Md. Kamrozzaman Bhuiyan, Manar D. Samad
The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.
By Kai Ming Ting, Wei-Jie Xu, Hang Zhang
arXiv:2607. 21088v1 Announce Type: new Abstract: Deep subspace clustering plays a critical role in applications involving multivariate spatiotemporal data, such as sea ice monitoring, disease spread analysis, and tracking neuro-degeneration over time.
By Francis Ndikum Nji, Vandana Janeja, Jianwu Wang
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.
By Joshua Dimasaka, Christian Gei{\ss}, Emily So
The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.
By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi
arXiv:2606. 08712v1 Announce Type: cross Abstract: Purpose: Spatial transcriptomics (ST) enables gene expression measurements within the tissue context.
By Hongyi Yu, Yaoyu Fang, Jiahe Qian, Xinkun Wang, Lee A. Cooper, Bo Zhou
The paper introduces federated soft clustering for devices in a federated learning network, each fitting a personalized Gaussian mixture model. It proposes Generalized Total Variation Minimization (GTVMin) to couple local maximum likelihood problems via a graph regularizer that penalizes discrepancies between connected nodes’ models. Three discrepancy measures are compared: a squared Euclidean distance requiring component matching, a Monte‑Carlo approximated Kullback‑Leibler divergence, and a closed‑form maximum mean discrepancy; all are optimized with synchronous projected gradient updates, with a convergence guarantee for the smooth MMD instance.
By Shamsiiat Abdurakhmanova, Alexander Jung