CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
Read the original on arXiv Machine Learning →CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.