arXiv Machine Learning

Addressing Imbalance in Multi-Label Data via Label-Specific Distance-based Oversampling

arXiv:2606. 05927v1 Announce Type: new Abstract: The complex imbalanced label distribution poses a crucial challenge to multi-label classification, as most classifiers are biased towards the majority class and high-frequent labels.

arXiv Machine Learning
Jul 30

The Advantage of Fine-Grained Training

arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.

By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Machine Learning
Jul 15

Calibratable Disambiguation Loss for Multi-Instance Partial-Label Learning

arXiv:2512. 17788v2 Announce Type: replace Abstract: Multi-instance partial-label learning (MIPL) is a weakly supervised framework that extends the principles of multi-instance learning (MIL) and partial-label learning (PLL) to address the challenges of inexact supervision in both instance and label spaces.

By Wei Tang, Yin-Fang Yang, Weijia Zhang, Min-Ling Zhang
arXiv Machine Learning
Aug 24

Benchmarking noisy label detection methods

arXiv:2510.16211v2 Announce Type: replace Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...

By Henrique Pickler, Jorge K. S. Kamassury, Danilo Silva
arXiv Machine Learning
Sep 3

A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research

The paper introduces a sampling algorithm based on a multivariate Bernoulli distribution to address challenges in multi‑label datasets where labels are non‑exclusive and vary widely in frequency. By estimating distribution parameters from observed label frequencies and computing weights for each label combination, the method produces weighted samples that reflect a target distribution while respecting label dependencies. Applied to Web of Science research articles labeled with 64 biomedical topics, the approach yielded a more balanced sub‑sample, improving representation of minority categories.

By Simon Chung, Colby J. Vorland, Donna L. Maney, Andrew W. Brown