arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.
By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv:2607. 06641v1 Announce Type: cross Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance.
By Felix Feldman, Joshua Harris, Timothy Laurence, Leo Loman, Ollie Higgins, Fan Grayson, Poonam Soma, Bethany Pace-Bonello, Michael Borowitz, Toby Nonnenmacher
arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.
By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
The paper introduces a method for generating multilingual reasoning traces for medical question answering using large language models. It creates 500,000 reasoning traces in English, Italian, and Spanish by retrieving medical information from Wikipedia and applies them to MedQA and MedMCQA datasets extended into Italian and Spanish. The approach improves performance in both few‑shot in‑context learning and supervised fine‑tuning, achieving state‑of‑the‑art results for 8B‑parameter LLMs and releasing all resources for further research.
By Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri
MMTClinic is a new benchmark that tests large language models on complex reasoning and question‑answering tasks involving clinical time‑series data. It combines text, medical images, and multivariate physiological signals to create 30,000 QA pairs—including 15,000 multiple‑choice and 15,000 open‑ended questions—in five languages (English, Hindi, Bengali, Marathi, and Tamil). The benchmark covers mortality prediction, heart‑rate forecasting, and SOFA score estimation, and evaluates 13 state‑of‑the‑art LLMs across zero‑shot, few‑shot, and chain‑of‑thought settings, revealing significant performance gaps across tasks, languages, and modalities.
By Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
arXiv:2608.22363v1 Announce Type: new
Abstract: Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on English, limitin...
By Jingbo Wang, Sendong Zhao, Haochun Wang, Bing Qin, Ting Liu
arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.
By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
arXiv:2608.22622v1 Announce Type: cross
Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficul...
By Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW. Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, Parisa Rashidi
arXiv:2607. 06452v1 Announce Type: cross Abstract: Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents.
By Taeyun Roh, Eunha Lee, Wonjune Jang, Sohyun Chung, Junha Jung, Jaewoo Kang
arXiv:2507. 02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering.
By Mohammad Anas Azeez, Rafiq Ali, Ebad Shabbir, Zohaib Hasan Siddiqui, Gautam Siddharth Kashyap, Jiechao Gao, Usman Naseem
arXiv:2609.06914v1 Announce Type: new
Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their a...
By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
Gaokerena is a family of compact Persian medical language models designed for consumer‑grade hardware. The first variant, Gaokerena‑V, was trained on a 90‑million‑token Persian medical corpus and 20,000 physician Q&A pairs, raising a translated medical MMLU benchmark score from 46.28% to 49.31%. A second variant, Gaokerena‑R, adds a Chain‑of‑Thought approach and two Reinforcement Learning with AI Feedback frameworks, achieving a higher benchmark score of 52.98% while also providing uncertainty estimates based on internal hidden states.
By Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami, Sadra Hakim, Zahra Kazemi, Amirhossein Poursina, Milad Tavakoli, Audrina Ebrahimi