The paper proposes a method for allocating a limited budget of expert annotations to optimize the accuracy of off-policy evaluation in settings where rewards are missing or noisy. By deriving variance‑optimal annotation probabilities for sequential, forward‑monotone protocols, the authors provide a batch‑adaptive implementation that can be applied to real data. Experiments on casenotes from a homelessness services nonprofit and on human‑preference votes from LMArena demonstrate substantial reductions in RMSE—up to 65% for housing placement and 68% for progress toward a housing application—when using only 40% or more of the full annotation budget.
By Woojin Chae, Ezinne Nwankwo, Haitong Qin, Angela Zhou
Offline reinforcement learning and off-policy evaluation evaluates dynamic treatment rules based on retrospectively collected data prior to deployment. In recent AI applications, state and reward info...
arXiv:2604. 14575v3 Announce Type: replace-cross Abstract: Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes.
By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv:2508. 13187v4 Announce Type: replace-cross Abstract: Homelessness is a persistent social challenge, impacting millions worldwide.
By Jonathan A. Karr Jr., Benjamin F. Herbst, Matthew L. Sisk, Xueyun Li, Ting Hua, Matthew Hauenstein, Georgina Curto, Nitesh V. Chawla
arXiv:2606. 17165v1 Announce Type: cross Abstract: Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost.
By Joel Persson, M{\aa}rten Schultzberg, Sebastian Ankargren
arXiv:2609.06294v1 Announce Type: new
Abstract: Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental s...
By Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez, Tamar Krishnamurti, Bryan Wilder
arXiv:2505. 17961v4 Announce Type: replace-cross Abstract: Causal inference typically assumes centralized access to individual-level data.
By R\'emi Khellaf, Aur\'elien Bellet, Julie Josse
The paper introduces a causal mediation framework to separate direct discrimination from structural inequality in AI-driven credit decisions. Using Pearl’s natural direct and indirect effects, it presents an identification strategy under treatment‑induced confounding and proposes a doubly‑robust estimator with efficiency guarantees. Empirical analysis of 89,465 mortgage applications shows that about 77% of racial denial disparities stem from financial mediators, while the remaining 23% represents a conservative lower bound on direct discrimination.
By Duraimurugan Rajamanickam
arXiv:2606. 30932v1 Announce Type: new Abstract: Two-sided marketplaces connect distinct user groups whose interests often conflict -- improving outcomes on one side could degrade the other side's experience.
By Yufei Wu, Zhen Yan
arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.
By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv:2608. 00657v1 Announce Type: cross Abstract: Causal inference usually concerns a scalar treatment, yet in many problems the treatment is unstructured: a text, an image, or a sequence of clinical decisions.
By Kevin Christian Wibisono, Yixin Wang
arXiv:2507. 14661v2 Announce Type: replace-cross Abstract: Semi-supervised domain adaptation (SSDA) seeks to achieve accurate predictions in a target domain with limited labeled target data by exploiting abundant source and unlabeled target data.
By Wooseok Ha, Yuansi Chen