arXiv AI

Multicalibration for Unbiased Model-Based Prevalence Estimation

arXiv AI
Aug 20

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

The paper introduces Debiased Inference with Multiple Imperfect Measurements (DMM), a framework that uses several error‑prone AI measurements to perform valid downstream statistical inference without requiring costly gold‑standard labels. By assuming conditional independence of the measurements given the true label and unit‑level features, DMM leverages CP decomposition and semiparametric theory to prove consistency and asymptotic normality of its estimator. Simulations demonstrate that DMM yields valid inference and can improve efficiency when additional imperfect measurements are available, and the authors provide diagnostics for the key independence assumption.

By Naoki Egami, Sooahn Shin
arXiv Machine Learning
Sep 17

Making Political Text Scaling Comparable: Infrastructure and Hyperparameter Sensitivity for 17 Algorithms

The paper argues that computational text‑based ideal point estimation (CT‑IPE) methods should be viewed as configurable measurement pipelines rather than fixed estimators. It presents a large‑scale comparative experiment involving 17 CT‑IPE algorithms, 5,537 runs, and about 4.25 million left‑right position estimates, and describes shared infrastructure that enables joint execution of these heterogeneous methods. Sensitivity analyses reveal that most algorithms exhibit low hyperparameter sensitivity (ICC < .10), with any remaining sensitivity concentrated in a few key researcher choices such as the language or embedding model, seed keyword lists, and number of topics.

By Patrick Parschan