arXiv Machine Learning

proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference

arXiv:2607. 24401v1 Announce Type: cross Abstract: Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly.

arXiv Machine Learning
Jul 31

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

arXiv:2607. 27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions.

By Aakash Bhagat, Shashank Choudhary
arXiv Machine Learning
Sep 23

From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI

The paper presents a layered framework for evaluating conversational AI by aligning offline proxy signals with online A/B experiment outcomes. It introduces a three‑step alignment chain—behavioral label to product outcome, classifier to candidate behavior, and offline signal to experiment effect—alongside an audit protocol that compares confidence intervals and rankings. In a real‑world deployment, the composite proxy achieved 81.1% F1 versus 34.3% for the raw classifier, correctly predicting direction on all 113 contrasts and enabling efficient prioritization of candidate models before costly online testing.

By Xuanyi Li, Vaskar Nath, Hossein Amirkhani, Jay Li, Alex Deng
arXiv Machine Learning
Aug 11

Demystifying Prediction Powered Inference

arXiv:2601. 20819v2 Announce Type: replace-cross Abstract: Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science.

By Yilin Song, Dan M. Kluger, Harsh Parikh, Tian Gu
arXiv Machine Learning
Aug 27

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

The paper evaluates the robustness of automated fact‑checking systems by cross‑benchmarking nine models—including random baselines, fine‑tuned transformers, zero‑shot LLMs, and top AVeriTeC 2025 systems—across four datasets from scientific, open‑web, and climate domains. It finds that fine‑tuned models outperform zero‑shot LLMs on ClimateCheck, that system rankings vary strongly with domain and metric, and that replacing retrieved evidence with gold annotations boosts veracity accuracy by 14–22 points, underscoring retrieval as the main bottleneck. The authors provide code, pre‑processed datasets, and results to enable reproducible research.

By Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck
arXiv Machine Learning
4d ago

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu