arXiv:2609.22161v1 Announce Type: cross
Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how...
By Yuzheng Fan, Haochun Wang, Sendong Zhao, Xiao Han, Ming Ma, Bing Qin
arXiv:2509. 08604v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-training or fine-tuning on medical data to enhance domain-specific accuracy and safety.
By Anran Li, Lingfei Qian, Mengmeng Du, Yu Yin, Yan Hu, Zihao Sun, Yihang Fu, Hyunjae Kim, Erica Stutz, Xuguang Ai, Qianqian Xie, Rui Zhu, Jimin Huang, Yifan Yang, Siru Liu, Yih-Chung Tham, Lucila Ohno-Machado, Hyunghoon Cho, Zhiyong Lu, Hua Xu, Qingyu Chen
arXiv:2609.22239v1 Announce Type: new
Abstract: Ambient AI is increasingly adopted in healthcare to automatically generate clinical notes from patient-clinician conversations, with the potential to s...
By Jakir Hossain, Yi-Fei Zhao, Hongjian Wang, Minmei Shih, Katie Leigh Mullen, Ahmad P. Tafti, Leming Zhou, Manoj Purohit, William Hogan, Jay Zeng, Elizabeth Skidmore, Yanshan Wang
arXiv:2607. 20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge, especially for smaller locally deployable models.
By Jessica Sena, Shesadree Priyadarshani, Miguel Contreras, Bharat Gandhi, Scott Siegel, Subhash Nerella, Parisa Rashidi
arXiv:2409. 07314v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows.
By Praveenkumar Kanithi, Cl\'ement Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan
The paper introduces CLEAR, an agentic framework designed to improve the reliability of large language models (LLMs) in medical contexts by adjudicating evidence from multiple sources. CLEAR generates candidate answers from three distinct pathways—parametric knowledge, locally curated corpora, and dynamically retrieved evidence—and then uses an aggregation verifier to evaluate agreement and conflict among these sources. An adjudication module decides whether to preserve or revise conclusions, employing override-guard and challenge-audit mechanisms, and initiates targeted follow-up searches when conflicts remain unresolved.
By Shuai Wang, Yize Zhao, Qingyu Chen
The paper investigates how large language models handle domain-specific jargon, comparing a general-purpose Llama‑3.1 with a version fine‑tuned on medical data. Two new medical jargon benchmarks reveal that the general model actually outperforms the fine‑tuned variant, and interpretability tools show the fine‑tuned model over‑emphasizes a few components linked to jargon predictions. Reweighting these components narrows the performance gap, and some jargon‑sensitive components also aid materials‑science tasks, indicating a partially domain‑agnostic representation of specialized terminology.
By Darin Keng, Zhewei Sun
The paper investigates how to incorporate biomedical knowledge graphs (KGs) into large language models (LLMs) for clinical diagnosis. It evaluates five KG task formulations, three training paradigms, two KGs, and three base LLMs, finding that all paradigms outperform a non‑finetuned baseline but differ in knowledge transfer behavior. Introducing Gradient Intervention Density (GID) and Gradient Distortion (GD) metrics, the study identifies a ‘surgical alignment’ regime—sparse, localized updates achieved by KG‑judgment training with KL regularization—that improves reasoning quality even when in‑domain accuracy is lower than task‑specific supervised fine‑tuning.
By Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao
arXiv:2508. 14817v2 Announce Type: replace-cross Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs).
By Skatje Myers, Dmitriy Dligach, Timothy A. Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew Churpek, Anoop Mayampurath, Majid Afshar
arXiv:2608.29582v1 Announce Type: cross
Abstract: Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigati...
By Yi Yu, Bo Wang, Chong Feng, Ge Shi, Xia Liu, Ziyi Yang, Xuewen Shi
arXiv:2607. 24838v1 Announce Type: cross Abstract: In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs).
By Seongwon Seo, Seung Hwan Cho, Young-Min Kim
arXiv:2509. 21530v2 Announce Type: replace Abstract: Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples.
By Dongkyu Cho, Miao Zhang, Rumi Chunara