Hugging Face Trending Papers

Discovering Latent Groups for Robust Classification

Machine learning models exploit spurious correlations, achieving high average accuracy but failing disproportionately on underrepresented subgroups. Existing methods address this by adjusting network parameters, guided either by subgroup annotations or inferred pseudo-group labels.

arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv Machine Learning
2d ago

Rank-Constrained Adaptation for Reliable Real-World Performance

arXiv:2602.06924v3 Announce Type: replace Abstract: Deep learning models trained to optimize average accuracy often exhibit systematic failures on particular subpopulations. In real-world settings li...

By Abinitha Gourabathina, Hyewon Jeong, Teya Bergamaschi, Marzyeh Ghassemi, Collin Stultz
arXiv Machine Learning
Sep 21

Robust Mixture Models for Algorithmic Fairness Under Latent Heterogeneity

The paper introduces ROME, a framework that learns latent group structure while optimizing worst-group predictive performance. ROME links latent-variable modeling with distributionally robust optimization through an Expectation-Maximization approach for linear models and a neural Mixture-of-Experts for nonlinear settings. Experiments on simulations and three real-world regression datasets show that ROME improves worst-group performance while maintaining competitive overall accuracy compared to existing group-aware and group-label-free robust learning methods.

By Siqi Li, Molei Liu, Yiwei Lyu, Ziye Tian, Chuan Hong, Nan Liu
arXiv AI
Aug 25

Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

The paper proposes a fairness-aware Mixture-of-Experts (MoE) framework that tackles routing-induced bias by applying subgroup reweighting to correct data imbalance and gate entropy regularization to prevent the gating network from collapsing onto subgroup attributes. This end-to-end approach keeps expert utilization balanced and interpretable, offering a clear view of how subgroups are allocated across experts. Experiments show that the method improves fairness while maintaining competitive predictive performance.

By Sunhee Hwang
arXiv Machine Learning
Sep 23

eXplaining to Learn (eX2L): Regularization Using Contrastive Visual Explanation Pairs for Distribution Shifts

The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.

By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez