arXiv Machine Learning By Leonardo Pellegrina, Fabio Vandin

Few-Shot Resampling for Scalable Statistically-Sound Data Mining

Read the original on arXiv Machine Learning →

arXiv:2606. 11235v1 Announce Type: new Abstract: A key step in knowledge discovery is the evaluation of data mining results.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 17

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.

By Trisha Mittal, Akshay Mehra, Joshua Kimball
arXiv Machine Learning
Sep 15

Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

The paper investigates how AI‑generated data, treated as anomalies linked to main data points, influences batch decompositions and undersampling in random datasets. Using redundancy graphs and iterative methods, it derives bounds for the minimum size of a strongly dissimilar decomposition and shows a phase transition: when anomalies are few, the minimum size depends on main data, but beyond a threshold it is dominated by anomalies. Additionally, the authors present a size criticality result for strong similarity in randomly undersampled datasets, illustrated with categorical data examples where the overall space far exceeds the dataset size.

By Ghurumuruhan Ganesan