arXiv AI

CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

arXiv:2607. 09880v1 Announce Type: cross Abstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support.

arXiv AI
Sep 7

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

MMTClinic is a new benchmark that tests large language models on complex reasoning and question‑answering tasks involving clinical time‑series data. It combines text, medical images, and multivariate physiological signals to create 30,000 QA pairs—including 15,000 multiple‑choice and 15,000 open‑ended questions—in five languages (English, Hindi, Bengali, Marathi, and Tamil). The benchmark covers mortality prediction, heart‑rate forecasting, and SOFA score estimation, and evaluates 13 state‑of‑the‑art LLMs across zero‑shot, few‑shot, and chain‑of‑thought settings, revealing significant performance gaps across tasks, languages, and modalities.

By Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
arXiv AI
Jun 16

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.

By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
arXiv AI
Sep 25

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a continuously maintainable benchmark for evaluating large language models (LLMs) in electronic health record (EHR) information retrieval. It presents a scalable framework that automatically generates question–answer pairs from longitudinal EHR notes, validated by nineteen clinicians. The benchmark allows assessment of multiple inference strategies and highlights that state‑of‑the‑art LLMs often miss clinically important information, especially when synthesis across documents is required.

By Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
Hugging Face Trending Papers
Sep 24

A Living Benchmark for Information Retrieval from Electronic Health Records

The paper introduces BRIE, a scalable framework that automatically creates question–answer pairs from longitudinal electronic health record notes, validated by nineteen clinicians. It offers a continuously maintainable benchmark for evaluating large language models in clinical settings, addressing limitations of manual, costly, and quickly outdated existing benchmarks. Experiments across nine LLMs and five inference strategies reveal that even state‑of‑the‑art systems often miss clinically important information, especially for synthesis‑heavy queries.

arXiv AI
Jun 8

TSAQA: Time Series Analysis Question And Answering Benchmark

arXiv:2601. 23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science.

By Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, Hanghang Tong
arXiv Machine Learning
Jul 20

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

arXiv:2607. 15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods.

By Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi
arXiv AI
Jun 9

From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG

arXiv:2603. 03292v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit high reasoning capacity in medical question-answering, but their tendency to produce hallucinations and outdated knowledge poses critical risks in healthcare fields.

By Wenhao Wu, Zhentao Tang, Yafu Li, Shixiong Kai, Mingxuan Yuan, Zhenhong Sun, Chunlin Chen, Zhi Wang