arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
arXiv:2608. 12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification.
By Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias
arXiv:2606. 24102v1 Announce Type: cross Abstract: Most electronic health record (EHR) foundation models encode clinical events as discrete event tokens from a fixed vocabulary and therefore cannot directly represent events containing unseen concepts or new combinations of concepts and attributes such as numeric values.
By Lin Lawrence Guo, Adam Paul Yan, Emily Vettese, Lillian Sung
arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.
By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv:2506. 04831v3 Announce Type: replace Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions, could support more proactive and personalized care, but requires modeling heterogeneous and longitudinal electronic health record (EHR) data.
By Chantal Pellegrini, Ege \"Ozsoy, David Bani-Harouni, Matthias Keicher, Nassir Navab
arXiv:2608. 16273v1 Announce Type: cross Abstract: Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research.
By Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
CT‑ΔBench is a new benchmark designed to evaluate vision‑language models on longitudinal 3D medical imaging difference reporting. It provides patient‑level split data, change‑aware metrics, and physician‑validated references to assess clinically meaningful interval changes between two CT scans. The paper also introduces DeltaMed, a baseline model that directly reasons over paired CT scans, and compares it to an indirect two‑stage approach that first generates single‑timepoint reports before differencing.
By Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang
arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.
By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.
By Florent Pollet, Matthew McDermott
arXiv:2607. 17508v1 Announce Type: cross Abstract: We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-specific interpretable models that synthesizes coefficient-space structure from natural-language task descriptions and a memory of previously learned task-specific predictors.
By Sazan Mahbub, Caleb Ellington, Zhiyuan Li, Yixin Yang, Souvik Kundu, Ben Lengerich, Eric P. Xing
arXiv:2605.11533v4 Announce Type: replace
Abstract: Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information,...
By Sike Xiang, Shuang Chen, Kevin Qinghong Lin, Jialin Yu, Yijia Sun, Philip Torr, Amir Atapour-Abarghouei
arXiv:2601.16753v2 Announce Type: replace-cross
Abstract: Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is...
By Xinyi Wang, Grazziela Figueredo, Ruizhe Li, Xin Chen