arXiv:2603.00842v2 Announce Type: replace
Abstract: Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remai...
By Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains a critical bottleneck for evidence-based medicine. While commercial black-box tools lack transparency, standard open-source RAG implementations frequently suffer from reasoning drift when handling complex, long-tail queries.
arXiv:2606. 29746v1 Announce Type: new Abstract: Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge bases, remains a critical bottleneck for evidence-based medicine.
By Maolin Liu, Fanyu Xu, Ruoqing Xu, Jiahang Zhang, Hao Wang, Rui Wang
MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.
By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendatio...
arXiv:2608. 08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows.
By Jack Stark, Srinath Saikrishnan, Vikram Seenivasan, Bernie Boscoe, Andrew Lizarraga, Tuan Do
arXiv:2608.24053v1 Announce Type: new
Abstract: Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space...
By Junjie Zhou, Ke Mei, Lei Li, Tianyi Wang, Fengyun Rao, Jing Lyu
arXiv:2606. 15231v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenarios.
By Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
arXiv:2609.21164v1 Announce Type: new
Abstract: Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decisio...
By Inyoung Choi, Sukwon Yun, Jiayi Xin, Jie Peng, Tianlong Chen, Qi Long
arXiv:2603.02790v2 Announce Type: replace
Abstract: Foundation models are changing the way we develop medical artificial intelligence. By learning broadly generalizable features across diverse data m...
By Michelle Stegeman (and on behalf of the UNICORN consortium), Lena Philipp (and on behalf of the UNICORN consortium), Fennie van der Graaf (and on behalf of the UNICORN consortium), Marina D'Amato (and on behalf of the UNICORN consortium), Cl\'ement Grisi (and on behalf of the UNICORN consortium), Luc Builtjes (and on behalf of the UNICORN consortium), Joeran S. Bosma (and on behalf of the UNICORN consortium), Judith Lefkes (and on behalf of the UNICORN consortium), Rianne A. Weber (and on behalf of the UNICORN consortium), James A. Meakin (and on behalf of the UNICORN consortium), Thomas Koopman (and on behalf of the UNICORN consortium), Anne Mickan (and on behalf of the UNICORN consortium), Mathias Prokop (and on behalf of the UNICORN consortium), Ewoud J. Smit (and on behalf of the UNICORN consortium), Fr\'ed\'erique Meeuwsen (and on behalf of the UNICORN consortium), Geert Litjens (and on behalf of the UNICORN consortium), Jeroen van der Laak (and on behalf of the UNICORN consortium), Bram van Ginneken (and on behalf of the UNICORN consortium), Maarten de Rooij (and on behalf of the UNICORN consortium), Henkjan Huisman (and on behalf of the UNICORN consortium), Colin Jacobs (and on behalf of the UNICORN consortium), Francesco Ciompi (and on behalf of the UNICORN consortium), Alessa Hering (and on behalf of the UNICORN consortium)
LatentVerse is a new framework that provides a web-based visual analytics platform and a command-line interface for analyzing multimodal latent representations. It unifies diagnostics for representation quality metrics and extends analysis to multimodal settings by decomposing embeddings into shared and modality-specific components. The authors evaluate the tool through simulations, real biomedical data analyses, and a user study, demonstrating its utility for interpretable evaluation of foundation model representations.
By Majd Alafrange, Samuel Friedman, John Kitonyo, Sana Tonekaboni, Mahnaz Maddah
The paper introduces ThoughtMed-1M, a large-scale medical visual question answering dataset built from de‑identified images and clinician‑generated commentaries, designed to capture structured clinical reasoning and image‑text alignment. Using this dataset, the authors train FOLTMed, a foundational large language model that achieves state‑of‑the‑art performance on 42 medical VQA benchmarks, with a macro accuracy of 85.4% and improved factuality and similarity metrics over existing models.
By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin