arXiv Statistics ML

A Causal Inference Approach for Evaluating Diagnostic Tests and AI-Enabled Medical Devices: From Effect Modification to Information-Augmented Decision-Making

arXiv:2608. 19501v1 Announce Type: cross Abstract: Diagnostic medical tests and devices provide useful information for evaluating the potential benefits and risks of therapeutic treatments.

arXiv AI
Sep 17

Information Set Emulation: Causal Certificates for AI Derived EHR Features

The paper introduces information set emulation, a method that attaches detailed causal certificates—such as source evidence, timing, and proposed causal roles—to AI‑derived features extracted from electronic health records (EHRs). These certificates provide auditable evidence for causal roles and guide whether a feature can be used for causal inference or should be routed to compatible reporting or separate analyses. The framework integrates with a joint EHR observation map and offers identification, estimation, and diagnostic tools under standard causal assumptions, illustrated through synthetic simulations and a finite‑world example.

By Takes Fujita (VRI), Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)
arXiv Machine Learning
Sep 24

Artificial intelligence surrogates for treatment effect estimation with before-and-after data

The paper proposes a method for estimating treatment effects using AI-generated predictions as surrogates, applied to paired before-and-after measurements for each treated individual. By comparing AI predictions before and after treatment, the approach can identify the average treatment effect on the treated under certain technical assumptions, even when clinical outcomes are never observed for treated subjects. When assumptions are questionable, the authors introduce prediction‑powered inference that corrects bias with a small set of observed outcomes, and validate the method with synthetic and cardio‑oncology data.

By Frances Dean, Anna Neufeld, Joshua Barrios, Geoffrey H Tison, Ahmed Alaa
arXiv AI
Aug 12

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

arXiv:2608. 10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions.

By Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
arXiv Computation and Language
Sep 23

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI: A Narrative Review and Conceptual Framework

The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.

By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv AI
Sep 25

An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer

The article introduces MedGPT Clinical Explorer (MCE), a conditional‑strategy framework designed to aid complex lung cancer decision‑making by explicitly mapping patient conditions to pathway eligibility, deferral, and redirection. In a study with 250 physicians across 98 institutions, MCE‑assisted strategies achieved higher Admissible Pathway Attainment Scores (APAS) than unaided or retrieval‑reference approaches, indicating more comprehensive inclusion of clinically relevant content and coherent links among pathways, conditions, and actions. The authors suggest that MCE’s shared decision object could improve transparency of omissions and contingencies, warranting prospective evaluation of its impact on workflow and patient outcomes.

By Daoyun Wang, Zhicheng Huang, Huaiyuan Sun, Jiaqi Xu, Xiaowei Xu, Zhibo Zheng, Zhongxing Bing, Yuxiao Lin, Yicheng Liang, Chao Gao, Bowen Xue, Kai Zhang, Song Xu, Wanpu Yan, Hui Xia, Lin Li, Xiang Yan, Mu Hu, Qianli Ma, Zhiqiang Xue, Xiaofang Liu, Zhihai Han, Nan Zhang, Chuanhao Tang, Tongmei Zhang, Lan Song, Zhaohui Zhu, Xuan Zeng, Shafei Wu, Hui Guan, Lei Deng, Huaxia Yang, Zeliang Lian, Wubin Sun, Yongxin Wang, Xiaohui Shen, Binlin Wang, Tiantian Gu, Yu Cui, Li Zhang, Shirui Wang, Naixin Liang
arXiv Machine Learning
Jul 31

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

arXiv:2607. 27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions.

By Aakash Bhagat, Shashank Choudhary
arXiv AI
Sep 17

Rethinking How We Evaluate Methodological Progress in Health AI

The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.

By Florent Pollet, Matthew McDermott