Large Language Models for Imbalanced Classification: Diversity makes the difference
arXiv:2510. 09783v2 Announce Type: replace-cross Abstract: Oversampling is one of the most widely used approaches for addressing imbalanced classification.
arXiv:2606. 05927v1 Announce Type: new Abstract: The complex imbalanced label distribution poses a crucial challenge to multi-label classification, as most classifiers are biased towards the majority class and high-frequent labels.
arXiv:2510. 09783v2 Announce Type: replace-cross Abstract: Oversampling is one of the most widely used approaches for addressing imbalanced classification.
arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
arXiv:2606. 07630v1 Announce Type: cross Abstract: Real-world datasets across image and text domains are often characterized by skewed class distributions and noisy annotations, which jointly degrade model performance, particularly on minority classes.
arXiv:2512. 17788v2 Announce Type: replace Abstract: Multi-instance partial-label learning (MIPL) is a weakly supervised framework that extends the principles of multi-instance learning (MIL) and partial-label learning (PLL) to address the challenges of inexact supervision in both instance and label spaces.
arXiv:2510.16211v2 Announce Type: replace Abstract: Label noise is a common problem in real-world datasets, affecting both model training and validation. Clean data are essential for achieving strong...
arXiv:2506. 10292v2 Announce Type: replace-cross Abstract: Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance on extensive labelled data.
The paper introduces a sampling algorithm based on a multivariate Bernoulli distribution to address challenges in multi‑label datasets where labels are non‑exclusive and vary widely in frequency. By estimating distribution parameters from observed label frequencies and computing weights for each label combination, the method produces weighted samples that reflect a target distribution while respecting label dependencies. Applied to Web of Science research articles labeled with 64 biomedical topics, the approach yielded a more balanced sub‑sample, improving representation of minority categories.
arXiv:2506. 01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.
arXiv:2609.26468v1 Announce Type: new Abstract: A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability o...
arXiv:2602. 08986v2 Announce Type: replace-cross Abstract: In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for more detailed or fine-grained classifications.
arXiv:2609.22734v1 Announce Type: cross Abstract: Clinical domain classification plays an important role in organizing and analyzing large volumes of unstructured medical text. However, medical trans...