The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.
By Yuya Shimizu
arXiv:2609.17238v1 Announce Type: cross
Abstract: High-dimensional data create challenges for causal effect estimation because identifying the covariates needed for correct model specification become...
By Muwon Kwon, Peter M. Steiner
arXiv:2609.36310v1 Announce Type: new
Abstract: Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the du...
By Ignacio Boero, Jonathan Nixon, Alejandro Ribeiro
arXiv:2602. 12972v2 Announce Type: replace-cross Abstract: In online advertising, marketing interventions such as coupons introduce significant confounding bias into Click-Through Rate (CTR) prediction.
By Siyun Yang, Shixiao Yang, Jian Wang, Di Fan, Kehe Cai, Haoyan Fu, Jiaming Zhang, Wenjin Wu, Peng Jiang
arXiv:2606. 02221v1 Announce Type: cross Abstract: Multi-task learning (MTL) aims to construct a joint model for multiple tasks by sharing a common representation across domains.
By Chengfeng Wu, Tao Zou, Yanru Wu, Jingge Wang
The paper introduces DCRMTA, an end‑to‑end framework for deep causal representation learning in multi‑touch attribution (MTA). It addresses a flaw in existing deconfounding pipelines that discard user‑related causal signals by explicitly preserving the causal impact of user features. Using structural causal modeling and adaptive counterfactual attention, DCRMTA produces invariant user representations and achieves up to a 5.2% relative improvement in PR‑AUC on real industrial datasets, while offering robust Shapley‑based credit allocations across marketing channels.
By Jiaming Tang, Jingxuan Wen, Liping Jing
arXiv:2607. 22313v1 Announce Type: cross Abstract: Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence.
By Masahiro Tanaka
arXiv:2604. 14575v3 Announce Type: replace-cross Abstract: Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes.
By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv:2610.00968v1 Announce Type: cross
Abstract: Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods requi...
By Arman Behnam, Binghui Wang
The paper introduces a control‑variable framework for deep neural networks to mitigate omitted variable bias, particularly shortcut learning where covariates like demographics influence predictions. It refits the final layer of a pre‑trained network using cross‑fitting with ridge penalisation, orthogonalises covariate effects, and marginalises predictions over covariate distributions to achieve unbiased, interpretable results. Experiments on simulated images and neuroimaging data show consistent estimation of true effects and performance close to models trained on unconfounded data.
By Manuel Pfeuffer, Roshan Prakash Rane, Kerstin Ritter, Sonja Greven
arXiv:2607. 10540v1 Announce Type: cross Abstract: We propose a two-stage estimator for structural mediation parameters that combines deep representation learning with G-estimation under the "no essential heterogeneity" (NEH) assumption.
By Roberto Faleh, Sofia Morelli, Holger Brandt
arXiv:2608. 10857v1 Announce Type: new Abstract: Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning.
By Viktoria Schuster, Sana Tonekaboni, Caroline Uhler