arXiv Machine Learning

Tabular Imbalanced Learning: A Survey, Benchmark, and Practical Guide

The paper surveys tabular imbalanced learning and introduces TILBench, a benchmark evaluating over 40 methods on 57 datasets. It presents a unified taxonomy of approaches and shows that no single method dominates across all settings, with performance depending on dataset regimes and computational constraints. Practical recommendations for method selection and future research directions are provided.

arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
arXiv AI
Jun 30

Beyond IID: How General Are Tabular Foundation Models, Really?

arXiv:2606. 30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry.

By Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzm\"uller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Ga\"el Varoquaux, Frank Hutter
arXiv Machine Learning
Sep 16

Collaborative Optimization of Multiclass Imbalanced Learning: Density-Aware and Region-Guided Boosting

The paper introduces a collaborative optimization Boosting model for multiclass imbalanced learning that integrates density and confidence factors to create a noise‑resistant weight update mechanism and a dynamic sampling strategy. The modules are tightly coupled to coordinate weight updates, sample region partitioning, and region‑guided sampling. Experiments on 40 public imbalanced datasets show the model significantly outperforms seven state‑of‑the‑art baselines.

By Chuantao Li, Zhi Li, Jiahao Xu, Jie Li, Sheng Li
arXiv Machine Learning
Sep 24

CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning

CRISP (Coreset Reduction via Importance-Stratified Pruning) is a linear-time method that reduces negative-class examples in highly imbalanced tabular datasets by allocating a budget across quantile strata of a proxy-model score and using sample weights to correct for unequal inclusion probabilities. On a production fraud dataset, CRISP cuts the training set from 25 M to about 1.70 M rows (a 93.2% reduction) while preserving 99.7% of the full-data Average Precision. In public benchmarks such as CriteoPrivateAds, CRISP consistently achieves the highest mean Average Precision across a range of majority reductions, with ablation studies highlighting budget allocation and inverse-propensity weighting as key contributors to its performance.

By Hardhik Mohanty, Indrayana Rustandi, Mohamadreza Sheibani