Sensitivity Sampling with Predictions for k-Means Clustering
arXiv:2607. 04949v1 Announce Type: new Abstract: We study the problem of k-means clustering on large datasets.
The paper introduces a general framework for multi-domain clustering using measure quantization, where a shared set of cluster prototypes is learned by minimizing a probability metric (e.g., Sinkhorn divergence or Maximum Mean Discrepancy) between each domain’s probability measure and the prototype measure. Data points are assigned to clusters either by nearest centroid or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini‑batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance, and experimental results on five multi‑domain benchmarks (image, audio, and sensor data) show that the Sinkhorn‑based method consistently outperforms classical and multi‑domain clustering baselines, even when scaling to hundreds of thousands of samples.
arXiv:2607. 04949v1 Announce Type: new Abstract: We study the problem of k-means clustering on large datasets.
The paper examines whether Deep Embedded Clustering (DEC) truly overcomes the fundamental limitations of k‑means clustering, such as handling clusters of arbitrary shapes, varied sizes, and densities. Through analysis, it finds that DEC does not exploit the underlying data distribution and therefore fails to address these limitations. Instead, a non‑deep learning approach that leverages distributional information of clusters can achieve the intended goals of deep clustering.
arXiv:2511. 17823v2 Announce Type: replace Abstract: Clustering algorithms have long been the topic of research, representing the more popular side of unsupervised learning.
arXiv:2607. 19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models.
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
arXiv:2607. 24237v1 Announce Type: new Abstract: Many existing clustering methods are designed based on a set-oriented definition---a cluster is a set of similar points---relying a point-to-point similarity function to find similar points.
arXiv:2607. 24405v1 Announce Type: new Abstract: In this work, we propose K-SurvMeans, a novel extension of K-Means for clustering survival data.
arXiv:2502. 08397v3 Announce Type: replace-cross Abstract: Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together.
arXiv:2606. 18833v1 Announce Type: new Abstract: This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection.
arXiv:2606. 10896v1 Announce Type: new Abstract: We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass.
The paper formalizes a geometric tradeoff between ambient separation and sampling gaps to determine when distinct manifold components can be reliably separated in clustering. It introduces a threshold phenomenon for mutual‑k‑nearest‑neighbor graphs, defining an uncertainty zone where the number of clusters cannot be identified. The authors propose Manifold‑Based Clustering (MBC), which outputs a bracket interval quantifying this uncertainty rather than forcing a single cluster count.
arXiv:2607. 15835v1 Announce Type: new Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids.