arXiv AI

Active Learning with Foundation Model Priors: Efficient Learning under Class Imbalance

arXiv:2606. 07630v1 Announce Type: cross Abstract: Real-world datasets across image and text domains are often characterized by skewed class distributions and noisy annotations, which jointly degrade model performance, particularly on minority classes.

arXiv Computation and Language
Aug 27

Loss-Based Active Learning for Neural Abstractive Summarization

Loss-Based Active Learning for Neural Abstractive Summarization proposes LOBSTER, an active learning framework that selects unlabeled documents similar to the model’s high‑loss training examples to correct specific weaknesses. The method is tailored for abstractive summarization, addressing instability and computational bottlenecks seen in prior work. Experiments on three benchmark datasets and two backbone models show that LOBSTER matches or surpasses state‑of‑the‑art performance while speeding up query selection by up to 665×.

By Michail Ioannou, Tatiana Passali, George Michalopoulos, Grigorios Tsoumakas
arXiv Machine Learning
Jul 21

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

arXiv:2607. 17653v1 Announce Type: cross Abstract: Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data.

By Jing Li, Pan Liu, Meng Zhao, Wanli Xue, Yanhong Yang, Xu Cheng, Fan Shi, Jianhua Zhang, Qinghua Hu, Shengyong Chen
arXiv Machine Learning
Sep 2

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.

By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv AI
Jun 2

ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning

arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.

By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause