Black-Box Knowledge Transfer across Distinct Feature Sets
arXiv:2608. 12403v1 Announce Type: cross Abstract: Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation.
arXiv:2412. 18081v3 Announce Type: replace-cross Abstract: We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets.
arXiv:2608. 12403v1 Announce Type: cross Abstract: Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation.
arXiv:2608. 20255v1 Announce Type: cross Abstract: This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups.
arXiv:2607. 03005v1 Announce Type: new Abstract: In high-dimensional Ising model estimation, target sample sizes are often limited, and effectively using auxiliary binary datasets of unknown relevance remains challenging.
arXiv:2605.04469v2 Announce Type: replace-cross Abstract: Large-scale population-level datasets, such as the UK Biobank and the All of Us Research Program, often lack covariates needed for a specific...
arXiv:2606. 08691v1 Announce Type: new Abstract: Modern data-driven applications increasingly involve learning from multiple heterogeneous sources, where a target dataset is limited but related information is available across domains.
Guided Adversarial Robust Transfer (GART) learning is a new transfer learning method that relaxes the requirement for source data to closely resemble the target population. By optimizing an adversarial loss over a mixture of source distributions, GART achieves faster convergence and improved prediction performance when target data are scarce. Experiments on simulated data and on multi‑institutional biobank‑linked electronic health records for high‑density lipoprotein cholesterol demonstrate higher robustness and accuracy compared to existing transfer learning approaches.
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
arXiv:2609.24086v1 Announce Type: cross Abstract: Large-scale multipurpose cohort studies and biobanks often omit covariates needed for specific downstream analyses. We study target-population infere...
arXiv:2507.23768v2 Announce Type: replace-cross Abstract: Existing methods for transfer learning struggle to deal with situations where the source datasets are limited and not guaranteed to be well-a...
arXiv:2606. 14023v1 Announce Type: cross Abstract: Optimal Transport has become recently a powerful method for domain adaptation by aligning source and target distributions.
The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.
arXiv:2303. 08777v3 Announce Type: replace-cross Abstract: Cross-validation is one of the most widely used tools for risk estimation and model selection in statistics and machine learning, yet its theoretical properties when embedded in a learning procedure remain insufficiently understood.