The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.
By Kai Ming Ting, Wei-Jie Xu, Hang Zhang
arXiv:2608. 00346v1 Announce Type: new Abstract: Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy.
By Pulock Das, Yina Hou, Md. Kamrozzaman Bhuiyan, Manar D. Samad
arXiv:2604. 07085v2 Announce Type: replace Abstract: In electronic health records (EHRs), clustering patients and distinguishing disease subtypes are key tasks to elucidate pathophysiology and aid clinical decision-making.
By Manar D. Samad, Yina Hou, Shrabani Ghosh
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
arXiv:2608.23182v1 Announce Type: cross
Abstract: We present a comparative study of label-free metrics for assessing the quality of representations in deep neural networks to understand their reliabi...
By Daniel Richards Arputharaj, Daniel J\"onsson, Gabriel Eilertsen
arXiv:2607. 06887v1 Announce Type: new Abstract: Most self-supervised image clustering models, actually almost all deep learning approaches, are based on gradient descent: In order to calculate the loss, every optimization step requires a clearly defined target, whether a contrastive split, a masked patch or entity, an EMA-teacher output, a pseudo-label, or a differentiable information-theoretic functional.
By Canlin Zhang, Xiuwen Liu
arXiv:2606. 06342v1 Announce Type: cross Abstract: Topological Data Analysis (TDA) offers a principled, intrinsic lens for comparing neural representations.
By Yan Wang, Tianyang Hu
The paper reviews 50 image augmentation and generation techniques, categorizing them into ten groups, and conducts a large‑scale empirical study to assess their effectiveness as test generators for embedding‑based image retrieval systems. Using Amazon Titan and OpenCLIP embeddings, the authors evaluate the techniques across four dimensions—embedding‑space similarity, embedding uncertainty, semantic realism, and retrieval failure rate—on CIFAR‑10, ImageNet‑1K, and an industrial dataset. Results show that weather simulation and SaSPA yield the highest uncertainty and failure rates while maintaining realistic visuals, whereas GAN‑based methods produce low realism due to synthetic artifacts.
By Yehan De Silva, Anirudh Sridhar, Armin Lotfy, Nafiseh Kahani, Yvan Labiche, Ziyu Wang, Frank Ouyang, Clare Carty, Azalia Shamsaei
arXiv:2606. 13007v1 Announce Type: cross Abstract: Clustering is fundamental to scRNA-seq analysis, serving as a cornerstone for identifying cell populations and resolving tissue heterogeneity.
By Ping Xu, Pengjiang Li, Tian Du, Zaitian Wang, Jiawei Gu, Ziyue Qiao, Pengfei Wang, Yuanchun Zhou
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2607. 22139v1 Announce Type: cross Abstract: Accurate pixel-level classification of coronary angiograms is critical for cardiovascular disease assessment, yet the field lacks standardized evaluation protocols.
By Dominik Bernard Lau, Hubert Malinowski, Jerzy Szyjut, Adam Brzeski, Tomasz Dziubich, Rados{\l}aw Targo\'nski, Tomasz Figatowski, Natalia Zieli\'nska
arXiv:2606. 07086v1 Announce Type: cross Abstract: Deep neural networks (DNNs) excel in computer vision tasks given large annotated datasets.
By Chen-Hsuan Fang, Wei-Hsinag Chen, Pin-Hsuan Yu, Jung-Hua Wang, Tsung-Wei Pan