arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.
By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
MMTClinic is a new benchmark that tests large language models on complex reasoning and question‑answering tasks involving clinical time‑series data. It combines text, medical images, and multivariate physiological signals to create 30,000 QA pairs—including 15,000 multiple‑choice and 15,000 open‑ended questions—in five languages (English, Hindi, Bengali, Marathi, and Tamil). The benchmark covers mortality prediction, heart‑rate forecasting, and SOFA score estimation, and evaluates 13 state‑of‑the‑art LLMs across zero‑shot, few‑shot, and chain‑of‑thought settings, revealing significant performance gaps across tasks, languages, and modalities.
By Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
arXiv:2507. 02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering.
By Mohammad Anas Azeez, Rafiq Ali, Ebad Shabbir, Zohaib Hasan Siddiqui, Gautam Siddharth Kashyap, Jiechao Gao, Usman Naseem
arXiv:2606. 13572v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios.
By Tanmoy Kanti Halder, Akash Ghosh, Subhadip Baidya, Arijit Roy, Sriparna Saha
Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and low-resource scenarios. This gap is critical in regions like rural India, where patients often express complex medical queries in native Indic languages and rely on multimodal inputs such as medical images.
PetQA is a Korean long‑form question‑answering benchmark designed to assess veterinary knowledge and clinical reasoning in large language and vision‑language models. It comprises 10,076 text‑only and 8,751 multimodal QA pairs about dogs and cats, with expert veterinarian answers, and a test split called PetQA‑Bench that includes question type and clinical condition annotations. The study evaluates 18 models across zero‑shot, retrieval‑augmented generation, and supervised fine‑tuning settings using ROUGE, BERTScore, and LLM‑as‑a‑judge metrics, revealing current models’ strengths and limitations and underscoring the need for better adaptation methods for clinically reliable veterinary AI; translated versions in five languages are also provided.
By Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park
arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.
By Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
arXiv:2512. 01241v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.
By David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh
DocTalkBN is a large-scale multimodal dataset of authentic expert telemedicine conversations in Bengali, comprising 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, and 10,274 host–doctor question–answer exchanges across 26 medical specialties. The dataset contains 1.7 million tokens and preserves the spontaneity and contextual richness of real medical interactions in a low-resource language. Three downstream tasks—medical triage classification, advice safety evaluation, and medical named entity recognition—are constructed to benchmark large language models and encoder-based baselines, demonstrating DocTalkBN’s practical usefulness for clinically grounded reasoning.
By Anik Saha, Fahmida Sultana Naznin, Sadatul Islam Sadi, Ananya Shahrin Promi, Wahid Al Azad Navid, Rifat Shahriyar
arXiv:2608. 10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust.
By Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
arXiv:2509. 02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behave in complex, high-stakes clinical scenarios.
By Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam
The paper introduces a method for generating multilingual reasoning traces for medical question answering using large language models. It creates 500,000 reasoning traces in English, Italian, and Spanish by retrieving medical information from Wikipedia and applies them to MedQA and MedMCQA datasets extended into Italian and Spanish. The approach improves performance in both few‑shot in‑context learning and supervised fine‑tuning, achieving state‑of‑the‑art results for 8B‑parameter LLMs and releasing all resources for further research.
By Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri