arXiv Machine Learning By Ghurumuruhan Ganesan

Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

Read the original on arXiv Machine Learning →

The paper investigates how AI‑generated data, treated as anomalies linked to main data points, influences batch decompositions and undersampling in random datasets. Using redundancy graphs and iterative methods, it derives bounds for the minimum size of a strongly dissimilar decomposition and shows a phase transition: when anomalies are few, the minimum size depends on main data, but beyond a threshold it is dominated by anomalies. Additionally, the authors present a size criticality result for strong similarity in randomly undersampled datasets, illustrated with categorical data examples where the overall space far exceeds the dataset size.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 17

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.

By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv Machine Learning
Sep 24

A Discrepancy-Based Perspective on Dataset Condensation

The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.

By Tong Chen, Raghavendra Selvan
arXiv AI
Sep 10

Revisiting Thinning Methods for Kernel Learning Problems

The paper introduces Backward Kernel Herding, an algorithm that iteratively removes data points to create representative subsets for kernel learning, achieving performance comparable to state‑of‑the‑art methods while speeding up subsampling when the reduced size is less than half the original dataset. It also proposes Flexible Kernel Thinning, an extension that allows construction of subsets of any size, not just successive halvings, and demonstrates that this method often yields the best predictive performance. Experiments on Gaussian Processes and Kernel Support Vector Machines show that Backward Kernel Herding excels in training‑time efficiency, while Flexible Kernel Thinning offers superior predictive accuracy and competitive memory usage, emphasizing the need to choose a reduction strategy based on the desired trade‑off between performance, cost, and memory.

By Blanca Cano-Camarero, Yago R. Aguado-Carrillo-de-Albornoz, \'Angela Fern\'andez-Pascual, Jos\'e R. Dorronsoro