arXiv Machine Learning

Artificial intelligence surrogates for treatment effect estimation with before-and-after data

The paper proposes a method for estimating treatment effects using AI-generated predictions as surrogates, applied to paired before-and-after measurements for each treated individual. By comparing AI predictions before and after treatment, the approach can identify the average treatment effect on the treated under certain technical assumptions, even when clinical outcomes are never observed for treated subjects. When assumptions are questionable, the authors introduce prediction‑powered inference that corrects bias with a small set of observed outcomes, and validate the method with synthetic and cardio‑oncology data.

arXiv Machine Learning
Jul 31

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

arXiv:2607. 27224v1 Announce Type: cross Abstract: External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions.

By Aakash Bhagat, Shashank Choudhary
arXiv AI
Sep 15

Causal multi-modal AI for personalized chemosensitivity prediction

A causal multi-modal AI model was developed to predict personalized chemosensitivity in breast cancer patients using routine pathology and clinical data. Trained on 9,141 patients from nine countries and validated on 1,994 patients from three countries, the model produced treatment-specific recurrence probabilities with near-perfect calibration and strong prognostic discrimination over 5- and 10-year horizons. It outperformed existing recurrence-score tests and could reduce chemotherapy prescriptions by 30% while maintaining recurrence-free rates, with predictive performance also transferring to non-breast cancers.

By Dhruva Biswas, Jeroen Berrevoets, Alec McClean, Linus Bao, Jungkyu Park, Ken G. Zeng, Joseph Cappadona, Cerise Tang, Chuwen Liu, Bartosz Machura, Yin Wu, Valerie Speirs, Hatem Soliman, Rohit Bhargava, Sheheryar Kabraji, Thaer Khoury, David Page, Brian Piening, Carlo Bifulco, Claudia Meurs, Pieter Westenend, Sylvie Chabaud, Jerome Lemonnier, Paul H. Cottu, Florence Dalenc, Fabrice Andre, Frederique Madeleine Penault-Llorca, Thomas Bachelot, Frederick Howard, Francisco J. Esteva, Kevin Kalinsky, Lajos Pusztai, Jan Witowski, Krzysztof J. Geras
arXiv AI
Sep 17

Rethinking How We Evaluate Methodological Progress in Health AI

The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.

By Florent Pollet, Matthew McDermott
Hugging Face Trending Papers
Aug 13

Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology

Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, medication changes, repeat interventions, and physiological measurements are recorded asynchronously and can change risk assessment over time.