arXiv:2610.01028v1 Announce Type: cross
Abstract: Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely...
By Sung Ho Jo, Seonghwi Kim, Wonsang Yun, Minwoo Chae
arXiv:2502. 18975v2 Announce Type: replace Abstract: Machine learning models are inherently bound to the distribution of the training data, often exploiting non-causal shortcuts.
By Martin Surner, Abdelmajid Khelil, Ludwig Bothmann
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv:2606. 02830v1 Announce Type: new Abstract: Real-world datasets often contain spurious correlations that are not causally related to the target label.
By Arda Fazla, Abolfazl Hashemi
arXiv:2607. 18278v1 Announce Type: cross Abstract: Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident despite being wrong.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2608.27704v1 Announce Type: new
Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
By Madhusudan Srinivasan, Namith Nishal Raphae