arXiv AI

Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data

arXiv:2601. 14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environments block patient-level data sharing and the persistent server connections required by iterative federated methods.

arXiv Machine Learning
Jun 24

Federated Survival Analysis in Healthcare: A Multi-Model Evaluation on Cross-Institutional Heterogeneous Breast Cancer Data

arXiv:2606. 23871v1 Announce Type: new Abstract: Survival analysis is central to clinical decision-making, yet reliable time-to-event models require large, diverse cohorts that are rarely available at a single institution, while privacy regulations restrict the centralization of patient data.

By Natalia Moreno-Blasco, Anusha Ihalapathirana, Pekka Siirtola, Miguel Fernandez-de-Retana
arXiv Machine Learning
Aug 5

Federated generative event models for tokenized electronic health records

arXiv:2608. 02939v1 Announce Type: new Abstract: Electronic health record foundation models are limited by institutionally siloed data and substantial performance degradation under cross-site transfer.

By Michael C. Burkhart, Luke Solo, Inhyeok Lee, S'Khaja Charles, Zewei "Whiskey" Liao, Kaveri Chhikara, Dema Therese, Wan-Ting Liao, Catherine A. Gao, William F. Parker, Brett K. Beaulieu-Jones
arXiv Statistics ML
Aug 25

Random Hazard Forests

Random Hazard Forests (RHF) is a survival tree ensemble that models how a patient's hazard changes over continuous time as new measurements arrive. RHF directly estimates a nonparametric hazard likelihood for predictable covariate processes, using an efficient working model to guide tree construction and then estimating flexible time‑varying hazards at each terminal node. By routing each tree based on the covariate state immediately before each time point, RHF can handle irregular and asynchronous covariate updates, and averaging across trees yields a pathwise hazard estimate that accurately captures changing risk in simulations and an intensive‑care application.

By Hemant Ishwaran, Eileen M. Hsich, Udaya B. Kogalur, Donald K. K. Lee
arXiv AI
Jun 24

A global log for medical AI

arXiv:2510. 04033v2 Announce Type: replace Abstract: Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure.

By Ayush Noori, Aaron E. Boussina, Hai Ho Bich, James Anibal, Julia Maslinski, Manuel Burger, Martin Faltys, Adam Rodman, Alan Karthikesalingam, Alessandro Blasimme, Annelia Itwaru, Ben Kaplan, Bilal A. Mateen, Christopher A. Longhurst, Daniel Yang, Dave deBronkart, Effy Vayena, Fedor Sergeev, Gauden Galea, Ha Thi Hai Duong, Harold F. Wolf III, Jacob Waxman, Joerg C. Schefold, Joshua C. Mandel, Juliana Rotich, Kenneth D. Mandl, Lily Poursoltan, Maryam Mustafa, Melissa Miles, Nigam H. Shah, Noa Dagan, Pavan Bodanki, Peter Lee, Philipp Koralus, Prathamesh Parchure, Prem Timsina, Ran D. Balicer, Robert Korom, Scott Mahoney, Seth Hain, Tien Yin Wong, Trevor Mundel, Vivek Natarajan, Ankit Sakhuja, Benjamin Glicksberg, C. Louise Thwaites, Gunnar R\"atsch, Karandeep Singh, David A. Clifton, Isaac S. Kohane, Marinka Zitnik
arXiv Computation and Language
Aug 25

Scaling Electronic Health Record Foundation Models for Population Health Management

The paper introduces Scaling Electronic Health Record Foundation Models for Population Health Management, a large‑scale model trained on billions of medical events from over 5 million patients in Taiwan and the United States. By aligning ICD codes across different health systems, the model achieves strong scaling and generalization across 11 chronic disease prediction tasks, outperforming tree‑based, general, and biomedical language models with high sensitivity at 99% specificity. It also demonstrates superior few‑shot performance on the EHRShot benchmark and shows that cross‑system alignment provides a stronger pretraining signal than single‑site duplication in data‑limited scenarios.

By Liwen Sun, Hao-Ren Yao, Ophir Frieder, Xiang Qian, Chenyan Xiong
arXiv Statistics ML
Sep 23

Efficient and scalable clustering of survival curves

arXiv:2512.16481v2 Announce Type: replace-cross Abstract: Survival analysis encompasses a broad range of methods for analyzing time-to-event data, with one key objective being the comparison of survi...

By Nora M. Villanueva, Marta Sestelo, Luis Meira-Machado
arXiv Machine Learning
Sep 16

Adaptive Bayesian Partner Selection for Federated Clinical Centers

Adaptive Bayesian Partner Selection (ABPS) is a peer‑to‑peer federated learning framework designed for heterogeneous clinical centers, where each center maintains a Beta‑Bernoulli posterior over prospective peers’ Shapley marginal utility and selects partners using an Upper Confidence Bound criterion. The lightweight propose‑reject mechanism allows centers to collaborate only when mutually beneficial, with the option to abstain from communication entirely. In experiments on 230 non‑IID ICU centers predicting in‑hospital mortality, the ABPS‑X variant achieves comparable accuracy to the strongest baseline (FedDyn) while reducing communication cost by 90% and enabling intentional isolation for many centers.

By Navid Seidi, Satyaki Roy, Sajal K. Das