MedGEN-Bench is a new benchmark for open‑ended multimodal medical generation that addresses limitations in current medical visual benchmarks, such as query‑image misalignment, closed‑ended answer spaces, and text‑centric outputs. The dataset contains 6,422 image‑text pairs across six imaging modalities, 15 clinical tasks, and 27 subtasks, including VQA, image editing, and contextual multimodal generation pairs. Evaluation combines reference‑based fidelity metrics with a structured, checklist‑guided assessment by a medical VLM judge, and preliminary results show that image‑output tasks remain unsaturated while contextual augmentation improves image‑instruction similarity.
By Junjie Yang, Yuhao Yan, Gang Wu, Rui Qian, Zhisheng Chen, Haijiang Li, Yuhe Wu, Qichao Zhao, Dawen Tian, Xiang Wan, Fenglei Fan, Wenjian Qin, Yongquan Zhang, Feiwei Qin, Changmiao Wang
arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv:2511.22232v2 Announce Type: replace-cross
Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical inter...
By Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen
arXiv:2608.22323v1 Announce Type: new
Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...
By Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.
By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.
arXiv:2409.16183v2 Announce Type: replace
Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...
By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.
By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
Lingshu is a medical‑specialized multimodal large language model that addresses key limitations of existing medical MLLMs, such as narrow knowledge coverage, hallucinations, and weak reasoning. The authors curate a comprehensive dataset combining medical imaging, texts, and general‑domain data, then train Lingshu in multiple stages to embed medical expertise and improve task performance. They also introduce MedEvalKit, a unified evaluation framework, and demonstrate that Lingshu outperforms current open‑source multimodal models on multimodal QA, text‑based QA, and medical report generation.
By Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, Yu Sun, Junao Shen, Chaojun Wang, Jie Tan, Deli Zhao, Tingyang Xu, Hao Zhang, Yu Rong
The paper introduces ThoughtMed-1M, a large-scale medical visual question answering dataset built from de‑identified images and clinician‑generated commentaries, designed to capture structured clinical reasoning and image‑text alignment. Using this dataset, the authors train FOLTMed, a foundational large language model that achieves state‑of‑the‑art performance on 42 medical VQA benchmarks, with a macro accuracy of 85.4% and improved factuality and similarity metrics over existing models.
By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin