arXiv Machine Learning

RCAP: Robust, Class-Aware, Probabilistic Dynamic Dataset Pruning

arXiv:2606. 11761v1 Announce Type: new Abstract: Dynamic data pruning techniques aim to reduce computational cost while minimizing information loss by periodically selecting representative subsets of input data during model training.

arXiv Machine Learning
Jul 3

Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning

arXiv:2607. 01474v1 Announce Type: new Abstract: Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard centralized techniques due to privacy and heterogeneity constraints.

By Haemin Park, Diego Klabjan, Martin W. Braun, Xiuqi Li, Balakrishnan Ananthanarayanan
arXiv Machine Learning
Sep 24

CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning

CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.

By Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani