The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.
By Yuya Shimizu
arXiv:2512. 10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data.
By Nick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith, Neel Nanda
arXiv:2604. 14575v3 Announce Type: replace-cross Abstract: Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes.
By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.
By Xianjie Guo, Yuwei Wang, Guodu Xiang, Xiaoli Tang, Kui Yu, Han Yu, Qiang Yang
The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.
By Naoki Egami, Sooahn Shin
arXiv:2606. 10607v1 Announce Type: cross Abstract: Causal discovery aims to uncover causal structures from observational data, which is crucial for real-world decision-making.
By Xinyu Li, Yuanyuan Wang, Haoxuan Li, Chuan Zhou, Erdun Gao, Bo Han, Tongliang Liu, Kun Zhang, Howard Bondell, Mingming Gong
arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.
By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
The paper introduces SGHA, a fully automated system that discovers research problems by structuring scientific literature into evidence-linked objects and a typed evidence graph. SGHA operates entirely on a local 9B open‑weight language model, avoiding proprietary frontier‑model APIs, and outputs traceable research‑problem families with assumptions, objectives, success criteria, and ambiguities. Comparative experiments in five machine‑learning domains show that SGHA’s corpus‑first, evidence‑constrained approach yields inspectable research‑problem formulation without relying on external models.
By Sarvesh Gharat, Junpei Komiyama
arXiv:2609.35845v1 Announce Type: cross
Abstract: Macroeconomic productivity metrics, such as Total Factor Productivity, register technological breakthroughs with multi-year reporting lags due to adm...
By Muhammad Sukri Bin Ramli
arXiv:2605. 05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences.
By Nicholas S. Kersting, Vittorio Castelli, Chieh Ting Yeh, Xinzhu Wang, Saad Taame, Khaoula Allak
arXiv:2409.14202v4 Announce Type: replace-cross
Abstract: The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process...
By Sukjin Han
TopiCLEAR is a framework that clusters document or sentence embeddings using adaptive dimensionality reduction to uncover low‑dimensional geometric structures that correspond to human‑interpretable topics. The method is evaluated on four benchmark datasets, showing strong agreement with human annotations, especially for short and informal texts. A Twitter case study demonstrates that TopiCLEAR yields more interpretable topics than LDA, recovering both annotated topic structure and coherent sub‑topics.
By Aoi Fujita, Taichi Yamamoto, Yuri Nakayama, Ryota Kobayashi