Optimal Transport Reweighting for Robust Learning under Spurious Correlations and Label Noise
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
arXiv:2606. 14965v1 Announce Type: new Abstract: Synthetic instance-dependent label noise (IDN) benchmarks are widely used to evaluate noisy-label learning methods, yet existing approaches typically generate noise through imperfect annotators or classifier raters, leaving the source of ambiguity implicit.
arXiv:2606. 11699v1 Announce Type: new Abstract: The performance of machine learning and deep learning models largely depends on the quality of the training data.
arXiv:2602.06924v3 Announce Type: replace Abstract: Deep learning models trained to optimize average accuracy often exhibit systematic failures on particular subpopulations. In real-world settings li...
arXiv:2502. 18975v2 Announce Type: replace Abstract: Machine learning models are inherently bound to the distribution of the training data, often exploiting non-causal shortcuts.
arXiv:2606. 02830v1 Announce Type: new Abstract: Real-world datasets often contain spurious correlations that are not causally related to the target label.