arXiv AI

PU classification under Non-SCAR: clustering-assisted logistic model with oversampling enhancement

arXiv Machine Learning
Sep 16

Collaborative Optimization of Multiclass Imbalanced Learning: Density-Aware and Region-Guided Boosting

The paper introduces a collaborative optimization Boosting model for multiclass imbalanced learning that integrates density and confidence factors to create a noise‑resistant weight update mechanism and a dynamic sampling strategy. The modules are tightly coupled to coordinate weight updates, sample region partitioning, and region‑guided sampling. Experiments on 40 public imbalanced datasets show the model significantly outperforms seven state‑of‑the‑art baselines.

By Chuantao Li, Zhi Li, Jiahao Xu, Jie Li, Sheng Li
arXiv Machine Learning
Sep 22

Focused PU learning from imbalanced data

arXiv:2605.14467v2 Announce Type: replace Abstract: We propose a new method of learning from positive and unlabeled (PU) examples in highly imbalanced datasets. Many real-world problems, such as dise...

By Elias Zavitsanos, Georgios Paliouras
arXiv Machine Learning
Sep 23

Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores

Density‑Ratio Rescoring (DRR) enhances a classifier trained with the original class prior by adding a survey‑raking dual score that reweights the majority class to match minority feature moments within a tolerance. The method standardizes both the dual and base scores, combines them with equal weight, and uses the fitted dual directly for prediction without resampling or refitting the base model. Experiments on 24 tabular benchmarks and a gene‑expression cohort show that DRR improves average precision over the standardized base on every dataset, with a mean gain of 0.034, and outperforms a shared‑dual raking‑and‑relabeling resampler on most datasets.

By Dongha Kim, Seunghwan Park
arXiv AI
2d ago

Partial AUC Maximization from Positive-unlabeled Data

arXiv:2610.00284v1 Announce Type: cross Abstract: The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarize...

By Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama, Hiroshi Takahashi, Kazuki Adachi, Yasuhiro Fujiwara
arXiv AI
Jul 21

Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods

arXiv:2505. 13518v3 Announce Type: replace-cross Abstract: Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing predictions toward the majority class and degrading classifier performance.

By Behnam Yousefimehr, Mehdi Ghatee, Javad Fazli, Shervin Ghaffari, Zahra Rafei, Mohammad Amin Seifi, Sajed Tavakoli, Abolfazl Nikahd, Mahdi Razi Gandomani, Alireza Orouji, Ramtin Mahmoudi Kashani, Sarina Heshmati, Negin Sadat Mousavi
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Machine Learning
Jun 16

Imbalanced Classification under Capacity Constraints

arXiv:2605. 03289v2 Announce Type: replace-cross Abstract: Detecting observations from a minority class under severe class imbalance is a central challenge in applications such as fraud detection, medical screening, and industrial quality control.

By Daniel Fraiman, Ricardo Fraiman