arXiv Machine Learning

The Hidden Cost of Resampling: How Imbalance Correction Degrades Probability Calibration in Tree Ensembles

arXiv:2606. 29720v1 Announce Type: new Abstract: Resampling methods such as SMOTE and random under/over-sampling are standard tools for class-imbalanced classification, almost always evaluated by minority-class accuracy or F1.

arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
Hugging Face Trending Papers
Jun 24

When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?

Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood. This paper develops a framework for characterizing when synthetic minority augmentation can improve threshold-integrated and threshold-optimized metrics, including AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold \(\F_1\) score.

arXiv Machine Learning
Sep 24

CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning

CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.

By Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani
arXiv Machine Learning
Jul 30

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.

By Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
arXiv Machine Learning
Sep 23

Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores

Density‑Ratio Rescoring (DRR) enhances a classifier trained with the original class prior by adding a survey‑raking dual score that reweights the majority class to match minority feature moments within a tolerance. The method standardizes both the dual and base scores, combines them with equal weight, and uses the fitted dual directly for prediction without resampling or refitting the base model. Experiments on 24 tabular benchmarks and a gene‑expression cohort show that DRR improves average precision over the standardized base on every dataset, with a mean gain of 0.034, and outperforms a shared‑dual raking‑and‑relabeling resampler on most datasets.

By Dongha Kim, Seunghwan Park
arXiv Machine Learning
Sep 3

RCProb: Probabilistic rule extraction from classification tree ensembles

RCProb is a probabilistic extension of rule extraction from tree ensembles that improves probability estimates by using smoothed atomic class-conditional evidence and a support‑adaptive mixture for final rule probabilities. Compared to RuleCOSI+, RCProb reduces median paired log‑loss by 71.9% for random forests and 62.5% for gradient boosting, while also decreasing the number of extracted rules by about 38% for both ensemble types. The method shows significant improvements in calibration metrics such as Confidence‑ECE and competitive native probability estimates, with further gains possible through post‑hoc calibration.

By Josue Obregon