arXiv Machine Learning

Multi-Domain Clustering via Measure Quantization

The paper introduces a general framework for multi-domain clustering using measure quantization, where a shared set of cluster prototypes is learned by minimizing a probability metric (e.g., Sinkhorn divergence or Maximum Mean Discrepancy) between each domain’s probability measure and the prototype measure. Data points are assigned to clusters either by nearest centroid or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini‑batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance, and experimental results on five multi‑domain benchmarks (image, audio, and sensor data) show that the Sinkhorn‑based method consistently outperforms classical and multi‑domain clustering baselines, even when scaling to hundreds of thousands of samples.

arXiv Machine Learning
Aug 26

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.

By Kai Ming Ting, Wei-Jie Xu, Hang Zhang
arXiv Machine Learning
Sep 17

Bracketing Uncertainty in Clustering Under the Manifold Hypothesis

The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.

By Savik Kinger, Luciano Dyballa, Steven W. Zucker
arXiv Machine Learning
Jul 20

Data-Native Global Optimization for Big Data K-means Clustering

arXiv:2607. 15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids.

By Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan