arXiv Machine Learning

Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data

arXiv:2608. 00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy.

arXiv Machine Learning
Aug 26

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.

By Kai Ming Ting, Wei-Jie Xu, Hang Zhang
arXiv Machine Learning
Jul 9

Converge to Surprise: Evolutionary Self-supervised Image Clustering

arXiv:2607. 06887v1 Announce Type: new Abstract: Most self-supervised image clustering models, actually almost all deep learning approaches, are based on gradient descent: In order to calculate the loss, every optimization step requires a clearly defined target, whether a contrastive split, a masked patch or entity, an EMA-teacher output, a pseudo-label, or a differentiable information-theoretic functional.

By Canlin Zhang, Xiuwen Liu
arXiv AI
Sep 21

Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data

The paper introduces FedDCN, a federated deep clustering network that jointly optimizes reconstruction and clustering losses for high‑dimensional, heterogeneous data. It addresses challenges of non‑IID client data by generating synthetic augmentations and applying geometric regularization to align latent spaces. Experiments show the method’s effectiveness under both IID and non‑IID settings, and the authors outline future research directions.

By Morris Stallmann, Charalampos S. Kouzinopoulos, Marcin Pietrasik, Anna Wilbik
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Machine Learning
Sep 25

Selective Inference for Deep Clustering in Latent Spaces

The paper introduces a selective inference framework tailored for deep clustering that uses a fixed pretrained encoder to map high‑dimensional data into a latent space before clustering. It addresses the complex selection bias arising from the nonlinear transformation and offers a computationally tractable method to perform valid statistical tests on cluster differences. Experiments on synthetic data show controlled Type I error and higher power compared to conservative baselines, while genomic case studies demonstrate the ability to uncover significant cluster differences while properly accounting for selection bias.

By Eina Mizui, Tomohiro Shiraishi, Shunichi Nishino, Ichiro Takeuchi