arXiv:2606. 07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization.
By Mahdi Alkaeed
arXiv:2606. 17474v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based, limiting their ability to reflect the sequential, uncertain, and interactive nature of real-world care.
By Jiahui Niu, Huizi Yu, Wenkong Wang, Guangxin Dai, Jingxian He, Xiang Li, Zhiying Liang, Xinxin Lin, Kent CY So, Bryan YP Yan, Yun Kwok Wing, Yanqiu Xing, Xin Ma, Lizhou Fan
arXiv:2602.11391v5 Announce Type: replace
Abstract: Objective: This study develops and validates a patient simulation framework that aligns with the National Institute of Standards and Technology AI...
By Md Tanvir Rouf Shawon, Mohammad Sabik Irbaz, Hadeel R. A. Elyazori, Keerti Reddy Resapu, Yili Lin, Vladimir Franzuela Cardenas, K. Pierre Eklou, Farrokh Alemi, Kevin Lybarger
arXiv:2601.09717v2 Announce Type: replace-cross
Abstract: Online medical consultations contain sensitive health information whose privacy implications depend not only on the entities mentioned but al...
By Yiwei Yan, Guanfeng Liu
arXiv:2605. 18937v2 Announce Type: replace Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights.
By Rory Sayres, Kejia Chen, Ayush Jain, Matthew Thompson, Jonathan Richina, Xiang Yin, Jimmy Hu, Fan Zhang, Bob Lou, Mike Sanchez, Ines Mezerreg, Meredith Schreier, Hamsa Subramaniam, I-Ching Lee, Yugang Jia, Daniel Mcduff, Yossi Matias, Avinatan Hassidim, Dale Webster, Yun Liu, Jackie Barr, Quang Duong
MedConceal is a new benchmark for evaluating medical dialogue systems on hidden‑concern reasoning under partial observability. It features 300 curated cases and 600 clinician‑LLM interactions, using an interactive patient simulator that hides latent concerns and tracks their revelation and resolution through theory‑grounded communication signals. The benchmark assesses both confirmation (surfacing hidden concerns) and intervention (addressing the primary concern), revealing that current models excel on different metrics while human clinicians still outperform them on intervention success.
By Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo