arXiv Machine Learning By Jacob Carlson

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

Read the original on arXiv Machine Learning →

arXiv:2511. 01680v4 Announce Type: replace-cross Abstract: Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
6d ago

Econometrics with Pre-Trained Embeddings for Unstructured Data

The paper examines the use of pre‑trained deep‑learning embeddings as covariates in economic analyses of unstructured data. It identifies two main challenges: the mismatch between training data/tasks of pre‑trained models and the target economic task, and the identification problem of the embedding function. The authors propose sufficient conditions—particularly a transferability criterion—to guarantee convergence, introduce a bootstrap test to assess transferability without re‑estimating embeddings, and apply the framework to various double‑machine‑learning settings, including an empirical study of labor‑supply elasticity on Amazon Mechanical Turk using job‑description embeddings.

By Yuya Shimizu
arXiv AI
Jun 24

A Survey on Federated Causal Discovery and Inference

arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.

By Xianjie Guo, Yuwei Wang, Guodu Xiang, Xiaoli Tang, Kui Yu, Han Yu, Qiang Yang
arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin