arXiv Computer Vision By Weiren Zhao, Yi Dong, Cheng Chen

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

Read the original on arXiv Computer Vision →

SynerMedGen is a unified framework that aligns medical multimodal understanding with generation tasks through task alignment. It introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers representations learned during understanding to medical image synthesis. The model achieves strong zero‑shot performance on 22 synthesis tasks and outperforms state‑of‑the‑art specialized and unified models when combined with generation training, supported by a new 1M‑sample SynerMed dataset.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 20

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

The paper introduces MedUAG, a unified medical multimodal model that supports both understanding and generation tasks. It presents MedUAGCorpus, the largest dataset of over 6 million instances across 14 imaging modalities, and MedUAGBench, a benchmark covering 12 diverse generation tasks with standardized protocols. Experiments show that MedUAG performs strongly across many medical understanding and generation tasks, setting a competitive baseline for future medical multimodal systems.

By Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu
arXiv AI
Jul 28

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.

By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
Hugging Face Trending Papers
Jul 27

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.