arXiv Statistics ML
6d ago

Econometrics with Pre-Trained Embeddings for Unstructured Data

The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.

By Yuya Shimizu
arXiv AI
Sep 25

DCRMTA: Deep Causal Representation Learning for Multi-Touch Attribution

The paper introduces DCRMTA, an end‑to‑end framework for deep causal representation learning in multi‑touch attribution (MTA). It addresses a flaw in existing deconfounding pipelines that discard user‑related causal signals by explicitly preserving the causal impact of user features. Using structural causal modeling and adaptive counterfactual attention, DCRMTA produces invariant user representations and achieves up to a 5.2% relative improvement in PR‑AUC on real industrial datasets, while offering robust Shapley‑based credit allocations across marketing channels.

By Jiaming Tang, Jingxuan Wen, Liping Jing