Pragmatic DML with AI-Learned Representations
arXiv:2610.01935v1 Announce Type: cross Abstract: Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We...
The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.
arXiv:2610.01935v1 Announce Type: cross Abstract: Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We...
arXiv:2412. 18081v3 Announce Type: replace-cross Abstract: We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets.
arXiv:2511. 01680v4 Announce Type: replace-cross Abstract: Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.
arXiv:2608. 12403v1 Announce Type: cross Abstract: Pre-trained black-box predictive functions encode knowledge distilled from massive datasets and extensive computation.
arXiv:2604. 14575v3 Announce Type: replace-cross Abstract: Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes.
While publicly available electricity market data presents a valuable resource for forecasting research, the field lacks established benchmark datasets for standardized comparison. As a result, many st...
arXiv:2512. 10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data.
The paper presents a comparative study of six deep learning models—state-space, MLP, RNN, and Transformer-based architectures—for cross-border electricity price forecasting using publicly available data. It focuses on generalization across markets and evaluates performance under low-data target-market conditions (zero-shot, one-shot, few-shot) with a standardized dataset for the Germany‑Luxembourg bidding zone in 2024. Results show that N‑HiTS and NBEATSx perform competitively in limited‑data scenarios, while transformer models achieve comparable accuracy but require more adaptation and tuning, and that careful feature selection and hyperparameter tuning improve performance.
arXiv:2602. 16061v2 Announce Type: replace-cross Abstract: Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNAR): users with stronger opinions are more likely to respond, so standard estimators are biased and the estimand is not identified without additional assumptions.
The paper introduces an assumption‑lean framework that uses AI‑generated measurements as weak shadow variables to identify and infer population quantities when data are missing not at random. Weak shadow variables are outcome‑informative proxies that are conditionally independent of missingness given the true outcome and covariates, and they do not need to predict missing outcomes accurately. The authors derive sharp bounds via linear programs and propose a localized penalized estimator with a subsampling algorithm for confidence intervals, demonstrating in semi‑synthetic experiments that the resulting intervals are substantially narrower and more accurate than classical MNAR methods.
arXiv:2505. 16319v3 Announce Type: replace Abstract: Accurate demand estimation is critical for the retail business in guiding the inventory and pricing policies of perishable products.
arXiv:2607. 22313v1 Announce Type: cross Abstract: Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence.