arXiv AI

Perspective on Bias in Biomedical AI: Preventing Downstream Healthcare Disparities

arXiv:2604. 14514v2 Announce Type: replace Abstract: Healthcare disparities persist across socioeconomic boundaries, often attributed to unequal access to screening, diagnostics, and therapeutics.

arXiv AI
Aug 11

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

arXiv:2505. 16941v4 Announce Type: replace-cross Abstract: Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability.

By Vincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang, Florent Pollet, Young Sang Choi, Xinzhuo Jiang, Yuta Kobayashi, Yanwei Li, Sara Matijevic, Karthik Natarajan, Shalmali Joshi
arXiv AI
Sep 7

A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap

The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.

By Michael Bouzinier, Dmitry Etin
arXiv AI
Sep 17

Rethinking How We Evaluate Methodological Progress in Health AI

The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.

By Florent Pollet, Matthew McDermott
arXiv AI
Aug 20

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.

By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv Computation and Language
Aug 25

Scaling Electronic Health Record Foundation Models for Population Health Management

The paper introduces Scaling Electronic Health Record Foundation Models for Population Health Management, a large‑scale model trained on billions of medical events from over 5 million patients in Taiwan and the United States. By aligning ICD codes across different health systems, the model achieves strong scaling and generalization across 11 chronic disease prediction tasks, outperforming tree‑based, general, and biomedical language models with high sensitivity at 99% specificity. It also demonstrates superior few‑shot performance on the EHRShot benchmark and shows that cross‑system alignment provides a stronger pretraining signal than single‑site duplication in data‑limited scenarios.

By Liwen Sun, Hao-Ren Yao, Ophir Frieder, Xiang Qian, Chenyan Xiong
arXiv AI
Sep 3

General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population

The paper introduces the General Demographic Pre-trained (GDP) model, a lightweight foundation model that learns representations from the two most common clinical attributes—age and sex. By optimizing encoding and visit‑reordering strategies, GDP embeddings are shown to improve predictive performance when concatenated with raw features across various disease and geographic cohorts. The model outperforms several state‑of‑the‑art tabular foundation models and tree‑based algorithms, demonstrating that enriched demographic embeddings can enhance classification tasks while remaining fully compatible with standard classifiers.

By Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang
arXiv Computer Vision
Aug 31

Medical Imaging AI Competitions Lack Fairness

Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.

By Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Ma\v{s}ka, Ma\"elys Solal, Arya Yazdan-Panah, Vilma Bozgo, \"Omer S\"umer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim R\"adsch, Nadim Hammoud, Fruzsina Moln\'ar-G\'abor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jim\'enez-S\'anchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein
arXiv Computation and Language
Sep 23

Quantitative Evidence Mining for Plausibility-Aware Biomedical AI: A Narrative Review and Conceptual Framework

The article proposes a framework called quantitative evidence mining to transform biomedical findings into structured, context-rich evidence units. It outlines core elements such as claim, measured entity, value, comparator, population, conditions, temporal context, uncertainty, provenance, validation, and expert review. The authors present an eight-stage reference architecture and emphasize that plausibility should remain multidimensional rather than collapsed into a single truth label, linking extraction to evidence synthesis for applications like clinical trials, biomarker research, and knowledge-graph construction.

By Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon, Martin Hofmann-Apitius, Marc Jacobs
arXiv AI
Jun 2

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.

By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu