arXiv:2610.00873v1 Announce Type: cross
Abstract: In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between...
By Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu
arXiv:2603. 15158v2 Announce Type: replace Abstract: Addressing the domain adaptation problem becomes more challenging when distribution shifts across domains stem from latent confounders that affect both covariates and outcomes.
By Zahra Rahiminasab, Reza Soumi, Arto Klami, Samuel Kaski
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
arXiv:2606. 22775v2 Announce Type: replace-cross Abstract: Distribution shift between training and deployment is a pervasive challenge for modern AI systems.
By Zhewen Hou, Tian Zheng
arXiv:2608. 13133v1 Announce Type: cross Abstract: Distributional shifts arise when the target deployment environment differs from the source environment that generated the training data.
By Zhiyi Li, Xiaojie Mao, Yunbei Xu, Ruohan Zhan
arXiv:2609.30886v1 Announce Type: new
Abstract: Conformal prediction can lose coverage when the data distribution changes after deployment. We study adaptation using labeled source data and unlabeled...
By Seungjin Choi
arXiv:2606. 04164v1 Announce Type: cross Abstract: Data samples used for training often differ from those encountered during fine-tuning and deployment, and while ML models show promise, their performance remains limited when only small annotated datasets are available.
By Sotirios Vavaroutas, Yu Yvonne Wu, Ali Etemad, Cecilia Mascolo
arXiv:2406. 13944v2 Announce Type: replace-cross Abstract: This paper establishes the generalization error of pooled min-$\ell_2$-norm interpolation in transfer learning, where data from diverse distributions are available.
By Yanke Song, Kenneth Gu, Sohom Bhattacharya, Pragya Sur
arXiv:2411.02771v3 Announce Type: replace-cross
Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. While c...
By Lars van der Laan, Alex Luedtke, Marco Carone
arXiv:2606. 26053v1 Announce Type: cross Abstract: Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood.
By Zhengchi Ma, Pengfei Lyu, Anru R. Zhang
arXiv:2609.24086v1 Announce Type: cross
Abstract: Large-scale multipurpose cohort studies and biobanks often omit covariates needed for specific downstream analyses. We study target-population infere...
By Huali Zhao (School of Mathematics and Statistics, Huazhong University of Science and Technology), Ke Deng (Department of Statistics and Data Science, Tsinghua University)
Synthetic data augmentation is widely used to mitigate class imbalance, but its theoretical effects on score-based classification remain poorly understood. This paper develops a framework for characterizing when synthetic minority augmentation can improve threshold-integrated and threshold-optimized metrics, including AUROC, AUPRC, best-threshold balanced accuracy, and best-threshold \(\F_1\) score.