arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.
By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
arXiv:2509. 07605v2 Announce Type: replace-cross Abstract: Class imbalance poses a significant challenge to supervised classification, particularly in critical domains like medical diagnostics and anomaly detection where minority class instances are rare.
By Ali Nawaz, Amir Ahmad, Shehroz S. Khan
arXiv:2606. 29720v1 Announce Type: new Abstract: Resampling methods such as SMOTE and random under/over-sampling are standard tools for class-imbalanced classification, almost always evaluated by minority-class accuracy or F1.
By Zewen Liu
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu
The paper surveys tabular imbalanced learning and introduces TILBench, a benchmark evaluating over 40 methods on 57 datasets. It presents a unified taxonomy of approaches and shows that no single method dominates across all settings, with performance depending on dataset regimes and computational constraints. Practical recommendations for method selection and future research directions are provided.
By Ruizhe Liu, Jiaqi Luo
The paper introduces DUA-D2C, a Dynamic Uncertainty-Aware Divide2Conquer method that improves overfitting remediation in deep learning. It refines the traditional Divide2Conquer approach by dynamically weighting subset models based on a composite score of accuracy and normalized prediction entropy, allowing the central model to learn more from generalizable and confident edge models. The authors provide theoretical justification, show reduced model variance, and demonstrate significant generalization gains across image, audio, and text benchmarks, even when combined with standard regularizers like Dropout.
By Md. Saiful Bari Siddiqui, Md Mohaiminul Islam, Md. Golam Rabiul Alam
arXiv:2605. 28021v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world and safety-critical scenarios, where test inputs may deviate from the training distribution and overconfident predictions on unknown samples can lead to unreliable decisions.
By Fengqiang Wan, Qing-Yuan Jiang, Fu Shen, Yang Yang
arXiv:2412. 16209v5 Announce Type: replace Abstract: When using machine learning for imbalanced binary classification problems, it is common to subsample the majority class to create a (more) balanced training dataset.
By Nathan Phelps, Daniel J. Lizotte, Douglas G. Woolford
RoBell-RVFL is a lightweight, quality‑aware generalized bell random vector functional link network designed to address class imbalance and noisy data in real‑world datasets. It uses a dual‑strategy sample‑level weighting: unit weights preserve minority class information, while a probability‑weighted generalized bell membership function suppresses noisy majority samples in a kernel‑induced feature space. Experiments on UCI and KEEL benchmarks, including tests with up to 40% label noise, show that RoBell‑RVFL consistently outperforms recent RVFL variants, demonstrating the importance of adaptive, quality‑aware sample weighting for robust learning.
By A. Rahaman, A. Quadir, M. Tanveer
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.
By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv:2606. 01221v1 Announce Type: cross Abstract: Imbalanced learning is a critical challenge in machine learning, where underrepresented target values can bias models and degrade prediction performance on rare but important cases.
By Shermin Shahbazi, Hossein Mohammadi, Mohsen Afsharchi