arXiv Computation and Language

Benchmarking Clinical Decision Pathway Adherence in Large Language Models

The paper introduces MEGA-CDP, a benchmark designed to evaluate medical large language models (LLMs) on their ability to generate clinical decision pathways (CDPs) that adhere to clinical practice guidelines. MEGA-CDP is built from 2,274 English and Chinese guidelines, producing 42,353 clinical cases with explicit reference CDPs, and supports both single-turn and multi-turn interactions. Experiments on 16 LLMs reveal that reliable guideline adherence remains difficult, underscoring the need for CDP-focused evaluation and the potential of MEGA-CDP to advance medical LLM performance.

arXiv AI
Jul 16

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

arXiv:2607. 13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving.

By Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler
arXiv Computation and Language
Aug 27

Retrieval-Augmented Agentic Rubric Generation for Reliable Medical Response Evaluation

The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.

By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv AI
Aug 25

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

The paper investigates how medical large language models (LLMs) may exhibit narrative anchoring bias when presented with the same clinical case in different patient voices. Using the NarrativeShield SDoH MedQA dataset, the authors evaluate three Qwen2.5 instruction‑tuned LLMs (1.5B, 3B, 7B) on 300 clinical cases, reporting metrics such as persona‑level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. The 7B model achieves the highest accuracy (56.33 %) and correct consistency (40.33 %), yet narrative sensitivity errors remain substantial (31.67 %).

By Ahnaf Atef Choudhury, Ramkrishna Saha
arXiv AI
Jul 10

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

arXiv:2607. 08602v1 Announce Type: new Abstract: Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality.

By Peng Cui, Jitao Wang, Siyan Xue, Yao Huang, Haoming Xia, Dong Li, Dengxiang Liu, Weilin Wang, Liping Liu, Leida Zhang, Yunfu Cui, Tao Peng, Daolin Ji, Haitao Zhao, Wei Zhang, Xiaojuan Wang, Weijie Ma, Zongren Ding, Jinlong Li, Yuan Ding, Jiajing Zhao, Zhiyu Chen, Chengkun Yang, Ziyue Huang, Jiaqi Liu, Fusheng Liu, Yang Zhou, Xiaojuan Wang, Zhongquan Sun, Shiyun Bao, Xiaojun Wang, Ming Yang, Guangxin Li, Bin Shu, Yong Liao, Hongxuan Li, Yao Tang, Shizhong Yang, Yongyi Zeng, Yufeng Yuan, Yinpeng Dong, Jihui Hao, Jun Zhu, Jiahong Dong