arXiv Machine Learning

Heterogeneous transfer learning for high-dimensional regression with feature mismatch

arXiv:2412. 18081v3 Announce Type: replace-cross Abstract: We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets.

arXiv Machine Learning
Sep 14

Guided Adversarial Robust Transfer Learning with Source Mixing

Guided Adversarial Robust Transfer (GART) learning is a new transfer learning method that relaxes the requirement for source data to closely resemble the target population. By optimizing an adversarial loss over a mixture of source distributions, GART achieves faster convergence and improved prediction performance when target data are scarce. Experiments on simulated data and on multi‑institutional biobank‑linked electronic health records for high‑density lipoprotein cholesterol demonstrate higher robustness and accuracy compared to existing transfer learning approaches.

By Xin Xiong, Zijian Guo, Tianxi Cai
arXiv Statistics ML
Sep 22

Doubly robust target inference for generalized linear regression with completely missing covariates

arXiv:2609.24086v1 Announce Type: cross Abstract: Large-scale multipurpose cohort studies and biobanks often omit covariates needed for specific downstream analyses. We study target-population infere...

By Huali Zhao (School of Mathematics and Statistics, Huazhong University of Science and Technology), Ke Deng (Department of Statistics and Data Science, Tsinghua University)
arXiv Statistics ML
6d ago

Econometrics with Pre-Trained Embeddings for Unstructured Data

The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.

By Yuya Shimizu
arXiv Machine Learning
Jul 7

Distribution-free Deviation Bounds and The Role of Domain Knowledge in Learning via Model Selection with Cross-validation Risk Estimation

arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.

By Diego Marcondes, Cl\'audia Peixoto