arXiv Machine Learning

A Discrepancy-Based Perspective on Dataset Condensation

The paper introduces a unified framework for dataset condensation (DC) that generalizes existing methods by using discrepancy measures to quantify the distance between probability distributions. It extends the traditional goal of DC—creating a small synthetic dataset that preserves generalization—to include additional objectives such as robustness and privacy. The framework positions DC as a formal approximation problem, broadening its applicability across different machine learning regimes.

arXiv Machine Learning
Jun 17

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.

By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv Machine Learning
2d ago

How Many Samples Are Enough for Learning Across Domains?

The paper investigates how many data samples per domain are needed for effective learning across multiple domains. It derives criteria from learning bounds that reveal an inverse linear relationship between the number of training domains and the required samples per domain, offering theoretical guidance for dataset adequacy and construction. The study also establishes a close link between in-domain learning and out-of-domain generalization through new generalization bounds.

By Hong Zheng
arXiv Machine Learning
Sep 21

Sparse Priors for Efficient Distribution Learning

arXiv:2609. 20883v1 Announce Type: new Abstract: Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal.

By Saumya Goyal, Barnab\'as P\'oczos
arXiv Machine Learning
Sep 15

Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

The paper investigates how AI‑generated data, treated as anomalies linked to main data points, influences batch decompositions and undersampling in random datasets. Using redundancy graphs and iterative methods, it derives bounds for the minimum size of a strongly dissimilar decomposition and shows a phase transition: when anomalies are few, the minimum size depends on main data, but beyond a threshold it is dominated by anomalies. Additionally, the authors present a size criticality result for strong similarity in randomly undersampled datasets, illustrated with categorical data examples where the overall space far exceeds the dataset size.

By Ghurumuruhan Ganesan
arXiv Machine Learning
Jun 9

LARP: Learner-Agnostic Robust Data Prefiltering

arXiv:2506. 20573v4 Announce Type: replace-cross Abstract: Public datasets, crucial for modern machine learning and statistical inference, often contain low-quality or contaminated samples that can harm model performance.

By Kristian Minchev, Dimitar I. Dimitrov, Nikola Konstantinov