arXiv AI

Domain-Adapted Small Language Models for Reliable Clinical Triage

arXiv AI
Aug 12

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

arXiv:2608. 10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs).

By Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
arXiv AI
2d ago

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

The study evaluates counterfactual bias in ten open‑source large language models (LLMs) for pediatric Emergency Severity Index (ESI) prediction. By creating paired clinical vignettes that differ only in demographic or socioeconomic variables, the authors measure shifts in acuity assignment, finding that counterfactual sensitivity varies widely across model families and sizes. A fine‑tuned Qwen2.5‑7B model exhibited the lowest sensitivity, while larger or medical‑domain models sometimes showed greater shifts, highlighting the need for fairness assessment before clinical deployment.

By Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
arXiv Computation and Language
Sep 22

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

The study evaluates large language models (LLMs) on sequential emergency department triage, where acuity labels are predicted from progressively longer nurse‑patient conversations. Six LLMs were tested at five checkpoints on simulated and physician‑authored dialogues, showing a decline from moderate‑to‑substantial agreement on full records to only fair‑to‑moderate agreement at each checkpoint. The models consistently anchor on chief complaint exchanges and fail to integrate later evidence, yielding low agreement with clinicians (QWK 0.295 vs. 0.887‑0.929) and concentrating predictions on ESI‑2 and ESI‑3. whyItMatters":"The findings reveal that LLMs, despite strong offline performance, cannot reliably handle the sequential nature of real‑time triage, highlighting a critical gap for safe deployment in emergency settings."

By Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan
arXiv AI
Jun 11

Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction

arXiv:2605. 04221v2 Announce Type: replace-cross Abstract: Clinical named entity recognition from dental progress notes is challenging because documentation is highly unstructured, domain-specific, and often privacy-sensitive.

By Yao-Shun Chuang, Tushti Mody, Uday Pratap Singh, Shirindokht Shiraz, Chun-Teh Lee, Ryan Brandon, Muhammad F Walji, Xiaoqian Jiang, Bunmi Tokede
arXiv AI
Sep 15

A primer on evaluation methods for large language models in healthcare

arXiv:2609.14819v1 Announce Type: cross Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and...

By Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries, Alyssa Pradhan, Timothy J Keyes, Bernardo C Bizzo, Keith J Dreyer, Sarah F Mercaldo, James M Hillis
arXiv AI
3d ago

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

The paper introduces a benchmark of over 6,000 clinical triage scenarios, 7,000 physician annotations, and 225,000 large language model (LLM) responses to assess how LLMs perform under realistic variations in clinical text. The study finds that LLMs tend to recommend unnecessary care more often than physicians, especially when the input text is perturbed, and that LLM recommendations are more sensitive to gender and tone changes than human recommendations. These findings underscore the importance of deployment‑oriented evaluations that reflect expert physician behavior.

By Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
arXiv AI
Aug 25

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

The paper investigates how medical large language models (LLMs) may exhibit narrative anchoring bias when presented with the same clinical case in different patient voices. Using the NarrativeShield SDoH MedQA dataset, the authors evaluate three Qwen2.5 instruction‑tuned LLMs (1.5B, 3B, 7B) on 300 clinical cases, reporting metrics such as persona‑level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. The 7B model achieves the highest accuracy (56.33 %) and correct consistency (40.33 %), yet narrative sensitivity errors remain substantial (31.67 %).

By Ahnaf Atef Choudhury, Ramkrishna Saha
arXiv Machine Learning
Sep 25

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

This study benchmarks transformer models for Bangla medical named entity recognition (NER), comparing BanglaBERT, multilingual BERT (mBERT), XLM‑RoBERTa, and GPT‑4o mini under zero‑shot and few‑shot prompting. Across a full test set of 3,179 samples, fine‑tuned XLM‑RoBERTa achieves a new state‑of‑the‑art F1‑score of 0.5959, while BanglaBERT lags with 0.4937, suggesting that domain diversity outweighs language specificity. The analysis shows high performance on Medicine and Specialist entities (F1 > 0.83) but lower accuracy on Symptoms (F1 0.4367), and demonstrates that fine‑tuned transformers outperform prompt‑only approaches by a factor of 3.76.

By Rakib Abdullah, Md. Maruful Islam Maruf
arXiv AI
Sep 4

MIRA: A Bilingual Benchmark for Medical Information Response Audit

MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.

By Mengyu Xu, Qiaoxin Yang, Qianqian Wang, Xiwei Dai, Weiyi Wu, Chongyang Gao
arXiv Computation and Language
Sep 10

Auditable Emergency Triage for Maternal and Newborn Care in India

arXiv:2609.09356v1 Announce Type: new Abstract: At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand sup...

By Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini, Neelima Devadas, Amrita K Prasen, Nikhil Nalin, Santhosh SJ, Sreeram Nurani Ramasubramanian, Muhammed Afeer K, Anubhav Arora
arXiv Computation and Language
Aug 28

Benchmarking Clinical Decision Pathway Adherence in Large Language Models

The paper introduces MEGA-CDP, a benchmark designed to evaluate medical large language models (LLMs) on their ability to generate clinical decision pathways (CDPs) that adhere to clinical practice guidelines. MEGA-CDP is built from 2,274 English and Chinese guidelines, producing 42,353 clinical cases with explicit reference CDPs, and supports both single-turn and multi-turn interactions. Experiments on 16 LLMs reveal that reliable guideline adherence remains difficult, underscoring the need for CDP-focused evaluation and the potential of MEGA-CDP to advance medical LLM performance.

By Nuo Chen, Xinyang Jiang, Zilong Wang, Zhifei Zhang, Xiaoye Qu, Jiajun Deng, Yulan Guo, Cairong Zhao