arXiv AI

DADIR: Density-Aware Data-level Imbalanced Regression Framework

arXiv:2607. 17178v1 Announce Type: cross Abstract: Imbalanced learning addresses predictive modeling problems with underrepresented regions of the data distribution.

arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
arXiv Machine Learning
Sep 16

Collaborative Optimization of Multiclass Imbalanced Learning: Density-Aware and Region-Guided Boosting

The paper introduces a collaborative optimization Boosting model for multiclass imbalanced learning that integrates density and confidence factors to create a noise‑resistant weight update mechanism and a dynamic sampling strategy. The modules are tightly coupled to coordinate weight updates, sample region partitioning, and region‑guided sampling. Experiments on 40 public imbalanced datasets show the model significantly outperforms seven state‑of‑the‑art baselines.

By Chuantao Li, Zhi Li, Jiahao Xu, Jie Li, Sheng Li
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Computer Vision
Sep 1

FairReL: Deepfake Detection using Fairness-Aware Representation Learning

FairReL is a fairness‑aware representation‑learning framework for deepfake detection that targets two subgroup‑sensitive components: multi‑scale spatial features and fine‑tuning‑induced residual features. It uses an SVD‑decomposed backbone to isolate residuals and introduces Group‑Conditional Wavelet Decorrelation (GCWD) and Subspace‑Localised Mean Alignment (SLMA) losses to suppress subgroup imbalance and align subgroup means. Experiments on FF++, Celeb‑DF, DFD, and DFDC show that FairReL improves unseen‑dataset AUC by 3.9% and reduces subgroup FPR disparity by 10.2% compared to the state‑of‑the‑art fairness‑aware detector.

By Xiaoman Lu, Jiaqi Li, Shuntian Zheng, Huiping Chen, Yu Guan
arXiv AI
Aug 25

Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion

The paper introduces a framework to tackle modality imbalance in multimodal learning by focusing on sample-level variations. It defines a Modality Gap metric to measure prediction discrepancies, models the resulting bimodal distribution with a Gaussian Mixture Model, and uses Bayesian probabilities for soft separation of balanced and imbalanced samples. A two‑stage training process—Warm‑up and Adaptive Training—reallocates loss weights based on the GMM, strengthening alignment for imbalanced samples while favoring fusion for balanced ones, and shows superior performance over existing baselines.

By Zhiwen Yu, Zhaocheng Liu, Xiaoqing Liu, Huanqiang Zeng, C. L. Philip Chen
arXiv Machine Learning
Jul 8

Imbalance-Robust and Sampling-Efficient Continuous Conditional GANs via Adaptive Vicinal Learning and Auxiliary Regularization

arXiv:2508. 01725v5 Announce Type: replace Abstract: Recent advances in continuous conditional generative modeling, including Continuous conditional Generative Adversarial Network (CcGAN) and Continuous Conditional Diffusion Model (CCDM), estimate high-dimensional data distributions conditioned on scalar regression labels such as angles, ages, or temperatures.

By Xin Ding, Yun Chen, Yongwei Wang, Kao Zhang, Sen Zhang, Peibei Cao, Xiangxue Wang