arXiv AI

PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

PetQA is a Korean long‑form question‑answering benchmark designed to assess veterinary knowledge and clinical reasoning in large language and vision‑language models. It comprises 10,076 text‑only and 8,751 multimodal QA pairs about dogs and cats, with expert veterinarian answers, and a test split called PetQA‑Bench that includes question type and clinical condition annotations. The study evaluates 18 models across zero‑shot, retrieval‑augmented generation, and supervised fine‑tuning settings using ROUGE, BERTScore, and LLM‑as‑a‑judge metrics, revealing current models’ strengths and limitations and underscoring the need for better adaptation methods for clinically reliable veterinary AI; translated versions in five languages are also provided.

arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv AI
Jul 9

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

arXiv:2607. 06641v1 Announce Type: cross Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance.

By Felix Feldman, Joshua Harris, Timothy Laurence, Leo Loman, Ollie Higgins, Fan Grayson, Poonam Soma, Bethany Pace-Bonello, Michael Borowitz, Toby Nonnenmacher
arXiv AI
Aug 18

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.

By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang
arXiv AI
Sep 2

Multilingual Medical Reasoning for Question Answering with Large Language Models

The paper introduces a method for generating multilingual reasoning traces for medical question answering using large language models. It creates 500,000 reasoning traces in English, Italian, and Spanish by retrieving medical information from Wikipedia and applies them to MedQA and MedMCQA datasets extended into Italian and Spanish. The approach improves performance in both few‑shot in‑context learning and supervised fine‑tuning, achieving state‑of‑the‑art results for 8B‑parameter LLMs and releasing all resources for further research.

By Pietro Ferrazzi, Aitor Soroa, Rodrigo Agerri
arXiv AI
6d ago

MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain

MMTClinic is a new benchmark that tests large language models on complex reasoning and question‑answering tasks involving clinical time‑series data. It combines text, medical images, and multivariate physiological signals to create 30,000 QA pairs—including 15,000 multiple‑choice and 15,000 open‑ended questions—in five languages (English, Hindi, Bengali, Marathi, and Tamil). The benchmark covers mortality prediction, heart‑rate forecasting, and SOFA score estimation, and evaluates 13 state‑of‑the‑art LLMs across zero‑shot, few‑shot, and chain‑of‑thought settings, revealing significant performance gaps across tasks, languages, and modalities.

By Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
arXiv AI
Jun 16

EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

arXiv:2606. 15735v1 Announce Type: cross Abstract: Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making.

By Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
arXiv AI
Aug 25

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

arXiv:2608.22622v1 Announce Type: cross Abstract: Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficul...

By Miguel Contreras, Scott Siegel, Subhash Nerella, Jessica Sena, Jiaqing Zhang, Heng Sun, Hruday Tej Akkaladevi, Peiyu Lu, Jordan Rosen, Sumit Kapoor, Sasank Desaraju, Grace R. Thompson, Jacob Purcell, Michael Petrauskis, Philip KW. Hong, Meghan Brennan, Sarah Chrabaszcz, Tierra Smith, Ronnie Ren, Michel S. Kabbash, Ceyhun Haziroglu, Rushi Patel, Gabriel Gomez, Charlotte Chaiklin, Randy Leung, Kenneth N. John, Whitman Wiggins, Philip Kayser, Vincent Bird, Maria Bruzzone, Tyler J. Loftus, Azra Bihorac, Parisa Rashidi
arXiv AI
3d ago

A visual large language foundational model for medical image recognition using clinician-oriented social media

arXiv:2609.06914v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their a...

By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
arXiv Computation and Language
Sep 4

Gaokerena: A Small Persian Medical Language Model Family

Gaokerena is a family of compact Persian medical language models designed for consumer‑grade hardware. The first variant, Gaokerena‑V, was trained on a 90‑million‑token Persian medical corpus and 20,000 physician Q&A pairs, raising a translated medical MMLU benchmark score from 46.28% to 49.31%. A second variant, Gaokerena‑R, adds a Chain‑of‑Thought approach and two Reinforcement Learning with AI Feedback frameworks, achieving a higher benchmark score of 52.98% while also providing uncertainty estimates based on internal hidden states.

By Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami, Sadra Hakim, Zahra Kazemi, Amirhossein Poursina, Milad Tavakoli, Audrina Ebrahimi