The paper introduces Expected‑Severity‑Risk (ESR), a new objective for selecting questions in proactive medical dialogue that prioritizes reducing the expected severity of diagnostic errors rather than merely uncertainty. ESR uses population statistics to marginalize over possible answers and distills its rankings into a prefix‑only language policy, enabling deployment without teacher‑side risk computation. Experiments on DDxPlus show ESR cuts high‑severity diagnostic misses by 29.5% and boosts accuracy while adding only 0.14 extra questions per dialogue.
By Chenxuan Li, Xinrong Chen, Luyan Zhang, Peidong Jia, Runfan Zheng, Zhongyu Zhao, Xuecheng Shang, Peixing Wan
arXiv:2609.24480v1 Announce Type: cross
Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the conv...
By Kalash Shah, Kunal Singh, Snehan J, Shreyas Singh
arXiv:2608. 20331v1 Announce Type: cross Abstract: Personalized interpretation of medical reports has emerged as an increasingly important need among patients.
By Shiao Xie, Siyu Chen, Jianwei Lv, Bo Yuan, Yujin Wang, Xiandong Li
arXiv:2607. 02983v1 Announce Type: new Abstract: Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information.
By Shengyi Hua, Kangzhe Hu, Conghui He, Xiaofan Zhang, Shaoting Zhang
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements.
arXiv:2609.24290v1 Announce Type: new
Abstract: Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confident...
By Yunxiang Li, Xixin Wu, Helen Meng
arXiv:2608.21721v1 Announce Type: new
Abstract: Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and wh...
By Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.
By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
arXiv:2607. 18999v1 Announce Type: cross Abstract: Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is sufficient.
By Guofeng Zhang, Yizeng Quan, Huaiyi Fang, Jianwei Lv, Jinyao Liu, Xunxu Duan, Lening An, Yu Ouyang, Junfeng Wang
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models proposes a new framework to improve medical reasoning in LLMs. It introduces two key conditions—Causal Sufficiency and Proximal Learnability—to curate high-quality training trajectories, using agreement-based self-verification and dynamic entropy bounds. Experiments on medical multimodal and text-only benchmarks show that CARE outperforms competitors, reducing incorrect reasoning and enhancing training stability.
By Yucheng Zhou, Peng Luo, Qianning Wang, Chengzhong Xu, Jianbing Shen
arXiv:2608.28599v1 Announce Type: new
Abstract: Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosi...
By Qi Peng, Yi Cai, Changmeng Zheng, Xin Wu, Jiayuan Xie, Qing Li
arXiv:2609.24799v1 Announce Type: new
Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generi...
By Yeji Kim, Mi-Young Kim, Randy Goebel