arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.
By Trisha Mittal, Akshay Mehra, Joshua Kimball
The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.
By Tong Chen, Raghavendra Selvan
arXiv:2609.06394v1 Announce Type: cross
Abstract: Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computat...
By Diptarka Chakraborty, Satyaki Mukherjee, Gaurav Vallabhdas Revankar, Hoang-Son Tran
arXiv:2606. 11235v1 Announce Type: new Abstract: A key step in knowledge discovery is the evaluation of data mining results.
By Leonardo Pellegrina, Fabio Vandin
The paper introduces Backward Kernel Herding, an algorithm that iteratively removes data points to create representative subsets for kernel learning, achieving performance comparable to state‑of‑the‑art methods while speeding up subsampling when the reduced size is less than half the original dataset. It also proposes Flexible Kernel Thinning, an extension that allows construction of subsets of any size, not just successive halvings, and demonstrates that this method often yields the best predictive performance. Experiments on Gaussian Processes and Kernel Support Vector Machines show that Backward Kernel Herding excels in training‑time efficiency, while Flexible Kernel Thinning offers superior predictive accuracy and competitive memory usage, emphasizing the need to choose a reduction strategy based on the desired trade‑off between performance, cost, and memory.
By Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, \'Angela Fern\'andez-Pascual, Jos\'e R. Dorronsoro
arXiv:2606. 01155v1 Announce Type: cross Abstract: Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not.
By Boqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal, Maurice van Keulen, Mykola Pechenizkiy, Elena Mocanu, Torsten Hoefler, Decebal Constantin Mocanu