arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2607. 19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models.
By Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer
arXiv:2606. 16045v1 Announce Type: new Abstract: In the data selection problem, the objective is to choose a small, representative subset of data that can be used to efficiently train a machine learning model.
By Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan, David P. Woodruff, Samson Zhou
arXiv:2502. 08397v3 Announce Type: replace-cross Abstract: Clustering is a fundamental technique in data analysis and machine learning, used to group similar data points together.
By Anna Livia Croella, Veronica Piccialli, Antonio M. Sudoso
arXiv:2607. 24237v1 Announce Type: new Abstract: Many existing clustering methods are designed based on a set-oriented definition---a cluster is a set of similar points---relying a point-to-point similarity function to find similar points.
By Kai Ming Ting, Kaifeng Zhang, Sanjay Chawla
arXiv:2609.07974v1 Announce Type: cross
Abstract: Fairness in clustering has attracted sustained research interest, motivated by the need to ensure equitable representation of protected groups in mac...
By Kangke Cheng, Guanlin Mo, Shihong Song, Hu Ding
arXiv:2606. 18833v1 Announce Type: new Abstract: This paper introduces a semi-supervised clustering framework grounded in the statistical duality between grouping principles and anomaly detection.
By Nassir Mohammad
arXiv:2606. 05230v1 Announce Type: cross Abstract: Selecting a clustering algorithm and its hyperparameters without labels is a common difficulty in engineering machine learning pipelines that work with unsupervised analysis of sensor, image, or process data.
By Mahdi Shamsi, Soosan Beheshti
arXiv:2411. 01576v3 Announce Type: replace Abstract: The explainable clustering problem was first posed by Moshkovitz et al.
By Maximilian Fleissner, Maedeh Zarvandi, Debarghya Ghoshdastidar
The paper presents new algorithms for fair k‑center clustering in Euclidean spaces, where a dataset is divided into groups and each group has a limit on the number of centers that can be chosen. A parameterized approximation algorithm achieves a 2.732 ratio, which is improved to 2.414 with exponential time in k. By integrating this into a one‑pass streaming framework, the authors obtain streaming approximations of 4.464 (improvable to 3.828) and a polynomial‑time streaming algorithm with a 4.732 ratio, further reduced to 4.42, surpassing previous state‑of‑the‑art results. Experiments confirm that these methods outperform existing approaches in clustering accuracy.
By Zeyu Lin, Chaoqi Jia, Longkun Guo, Chao Chen
The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.
By Tong Chen, Raghavendra Selvan
arXiv:2609. 30477v1 Announce Type: cross Abstract: Exact Euclidean \(K\)-means partitions \(n\) observations into \(K\) unlabelled clusters, but the unrestricted search is generally exponential.
By Yordan P. Raykov, Max A. Little