arXiv Machine Learning

Sensitivity Sampling with Predictions for k-Means Clustering

arXiv:2607. 04949v1 Announce Type: new Abstract: We study the problem of k-means clustering on large datasets.

arXiv Machine Learning
Jun 16

Active Learning with Low-Rank Structure for Data Selection

arXiv:2606. 16045v1 Announce Type: new Abstract: In the data selection problem, the objective is to choose a small, representative subset of data that can be used to efficiently train a machine learning model.

By Vincent Cohen-Addad, Sasidhar Kunapuli, Vahab Mirrokni, Mahdi Nikdan, David P. Woodruff, Samson Zhou
arXiv Machine Learning
Sep 10

A Sub-4 Approximation for Fair $k$-Means

arXiv:2609.07974v1 Announce Type: cross Abstract: Fairness in clustering has attracted sustained research interest, motivated by the need to ensure equitable representation of protected groups in mac...

By Kangke Cheng, Guanlin Mo, Shihong Song, Hu Ding
arXiv AI
Sep 10

Parameterized and Streaming Algorithms for Euclidean Fair $k$-Center Clustering

The paper presents new algorithms for fair k‑center clustering in Euclidean spaces, where a dataset is divided into groups and each group has a limit on the number of centers that can be chosen. A parameterized approximation algorithm achieves a 2.732 ratio, which is improved to 2.414 with exponential time in k. By integrating this into a one‑pass streaming framework, the authors obtain streaming approximations of 4.464 (improvable to 3.828) and a polynomial‑time streaming algorithm with a 4.732 ratio, further reduced to 4.42, surpassing previous state‑of‑the‑art results. Experiments confirm that these methods outperform existing approaches in clustering accuracy.

By Zeyu Lin, Chaoqi Jia, Longkun Guo, Chao Chen
arXiv Machine Learning
Sep 24

A Discrepancy-Based Perspective on Dataset Condensation

The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.

By Tong Chen, Raghavendra Selvan