arXiv AI

Large Language Models for Imbalanced Classification: Diversity makes the difference

arXiv:2510. 09783v2 Announce Type: replace-cross Abstract: Oversampling is one of the most widely used approaches for addressing imbalanced classification.

arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Machine Learning
Sep 14

Bias-Corrected Data Synthesis for Imbalanced Learning

The paper introduces a bias‑correction method for synthetic oversampling in imbalanced learning. It estimates the loss discrepancy caused by the data generator using a held‑out majority subset and transfers this correction to the minority class under a uniform bias‑transfer assumption. The authors provide finite‑sample bounds for bias transfer and excess balanced risk, identify when SMOTE introduces significant bias, and demonstrate the method’s applicability to multi‑task learning and propensity‑score estimation, with empirical results showing greatest benefit when synthetic distortion is large.

By Pengfei Lyu, Zhengchi Ma, Linjun Zhang, Anru R. Zhang