Multi-Domain Clustering via Measure Quantization
Read the original on arXiv Machine Learning →The paper introduces a general framework for multi-domain clustering using measure quantization, where a shared set of cluster prototypes is learned by minimizing a probability metric (e.g., Sinkhorn divergence or Maximum Mean Discrepancy) between each domain’s probability measure and the prototype measure. Data points are assigned to clusters either by nearest centroid or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini‑batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance, and experimental results on five multi‑domain benchmarks (image, audio, and sensor data) show that the Sinkhorn‑based method consistently outperforms classical and multi‑domain clustering baselines, even when scaling to hundreds of thousands of samples.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.