arXiv:2603. 15158v2 Announce Type: replace Abstract: Addressing the domain adaptation problem becomes more challenging when distribution shifts across domains stem from latent confounders that affect both covariates and outcomes.
By Zahra Rahiminasab, Reza Soumi, Arto Klami, Samuel Kaski
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
arXiv:2606. 22775v2 Announce Type: replace-cross Abstract: Distribution shift between training and deployment is a pervasive challenge for modern AI systems.
By Zhewen Hou, Tian Zheng
arXiv:2608. 13133v1 Announce Type: cross Abstract: Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data.
By Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan
arXiv:2606. 04164v1 Announce Type: cross Abstract: Data samples used for training often differ from those encountered during fine-tuning and deployment, and while ML models show promise, their performance remains limited when only small annotated datasets are available.
By Sotirios Vavaroutas, Yu Yvonne Wu, Ali Etemad, Cecilia Mascolo
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
By Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur
arXiv:2606. 26053v1 Announce Type: cross Abstract: Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood.
By Zhengchi Ma, Pengfei Lyu, Anru R. Zhang
Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood. This paper develops a framework for characterizing when synthetic minority augmentation can improve threshold-integrated and threshold-optimized metrics, including AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold \(\F_1\) score.
arXiv:2608. 15783v1 Announce Type: cross Abstract: In transfer-learning settings, a model derived from abundant surrogate labels may be deployed in a target population where gold-standard outcomes are unobserved.
By Longtian Shi, Molei Liu, Doudou Zhou
arXiv:2209. 01754v5 Announce Type: replace-cross Abstract: The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed.
By Roshni Sahoo, Lihua Lei, Stefan Wager
arXiv:2502. 18049v5 Announce Type: replace-cross Abstract: Recent studies identified an intriguing phenomenon in recursive generative model training known as model collapse, where models trained on data generated by previous models exhibit severe performance degradation.
By Hengzhi He, Shirong Xu, Guang Cheng
arXiv:2511. 18945v4 Announce Type: replace Abstract: We propose a fully data-driven approach to designing mutual information (MI) estimators.
By German Gritsai, Megan Richards, Maxime M\'eloux, Kyunghyun Cho, Maxime Peyrard