arXiv AI
Sep 10

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

The study compared the performance of the AI system Doctorina, eight physicians, and four frontier language models on 150 synthetic Polish primary‑care consultations. Doctorina achieved 82.0% top‑1 concordance, outperforming physicians (57.0%) by 25 percentage points, and also led in diagnostic workup and treatment scores. A second run confirmed these advantages across all measured outcomes.

By Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou, Aleksey Ropan, Pavel Satalkin
arXiv Computation and Language
Aug 28

Benchmarking Clinical Decision Pathway Adherence in Large Language Models

The paper introduces MEGA-CDP, a benchmark designed to evaluate medical large language models (LLMs) on their ability to generate clinical decision pathways (CDPs) that adhere to clinical practice guidelines. MEGA-CDP is built from 2,274 English and Chinese guidelines, producing 42,353 clinical cases with explicit reference CDPs, and supports both single-turn and multi-turn interactions. Experiments on 16 LLMs reveal that reliable guideline adherence remains difficult, underscoring the need for CDP-focused evaluation and the potential of MEGA-CDP to advance medical LLM performance.

By Nuo Chen, Xinyang Jiang, Zilong Wang, Zhifei Zhang, Xiaoye Qu, Jiajun Deng, Yulan Guo, Cairong Zhao
arXiv AI
Jun 17

First, do NOHARM: towards clinically safe large language models

arXiv:2512. 01241v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.

By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
arXiv AI
Sep 15

A primer on evaluation methods for large language models in healthcare

arXiv:2609.14819v1 Announce Type: cross Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and...

By Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries, Alyssa Pradhan, Timothy J Keyes, Bernardo C Bizzo, Keith J Dreyer, Sarah F Mercaldo, James M Hillis
arXiv Computation and Language
Sep 23

Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation

The paper introduces a four‑stage framework—SFT, PG‑CoT, Dynamic, and K‑RL—to improve large language models for Traditional Chinese Medicine prescription generation. It addresses three key gaps: lack of auditable reasoning (SR Gap), failure to adjust prescriptions over time (LA Gap), and non‑enforcement of absolute contraindication rules (SC Gap). Experiments on 12 fine‑tuned models and 6 zero‑shot baselines show that the framework, particularly a 7B Mistral model, outperforms zero‑shot GPT‑5 on all three TCM evaluation metrics.

By Zheng Chen, ZhiCheng Du, Haoxuan Li, Peiwu Qin