arXiv Machine Learning By Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante

Multi-Domain Clustering via Measure Quantization

Read the original on arXiv Machine Learning →

The paper introduces a general framework for multi-domain clustering using measure quantization, where a shared set of cluster prototypes is learned by minimizing a probability metric (e.g., Sinkhorn divergence or Maximum Mean Discrepancy) between each domain’s probability measure and the prototype measure. Data points are assigned to clusters either by nearest centroid or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini‑batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance, and experimental results on five multi‑domain benchmarks (image, audio, and sensor data) show that the Sinkhorn‑based method consistently outperforms classical and multi‑domain clustering baselines, even when scaling to hundreds of thousands of samples.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 26

How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning

The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.

By Kai Ming Ting, Wei-Jie Xu, Hang Zhang