arXiv Machine Learning

Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.

Hugging Face Trending Papers
Jun 24

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these issues, we propose a tailored VLP framework featuring three key components: (1) a CNN-ViT hybrid encoder that replaces ViT's patch embedding with a 3D CNN backbone to efficiently capture local anatomical details while preserving global attention and compatibility with pre-trained cross-modal priors; (2) a disease-level contrastive learning mechanism using learnable query tokens to dynamically extract disease-specific semantics from full reports and align them with corresponding visual features, thereby disentangling distinct diseases within the same anatomical region; and (3) a diagnosis-aware prompt strategy that employs real clinical phrases and aggregated disease prototypes to bridge the pre-training-inference gap and enhance zero-shot diagnostic reliability.

arXiv Machine Learning
Jun 3

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

arXiv:2606. 03180v1 Announce Type: cross Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows.

By Jonggwon Park, Seongeun Lee, Junhyun Park, Hannah Yun, Hyunwoong Kim, Sohyun Jeong, Hyewon Kang, Byungmu Yoon, Kyoyun Choi
arXiv Computer Vision
Sep 22

MTMed3D: A Multi-Task Transformer-Based Model for 3D Medical Imaging

MTMed3D is a multi-task Transformer-based model that jointly performs 3D detection, segmentation, and classification in medical imaging. It uses a shared Transformer encoder to produce multi-scale features, with separate CNN decoders for each task. Evaluated on BraTS 2018 and 2019, it achieves strong results, especially in detection, while reducing computational cost and inference time compared to single-task models.

By Fan Li, Arun Iyengar, Lanyu Xu
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv AI
Jul 28

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.

By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv Computer Vision
Sep 2

Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT

Instance-Guided Report Anchoring (IGRA) is a model-agnostic module that links each abnormality instance in a chest CT to the corresponding finding in a radiology report during training, while discarding text components at inference. By reformulating free-text grounding as multi-label volumetric segmentation, IGRA allows all abnormality categories to be predicted in a single image-only forward pass. The method improves Dice scores by 22.5% over the strongest image-only baseline and matches state‑of‑the‑art performance on single-finding subsets, with consistent gains across multiple 3D segmentation backbones and datasets.

By Zhenyu Bu, Haoyan Ding, Chushu Shen, Xinyuan Zheng, Peiyu Duan, Xueqi Guo, Sepehr Farhand, Yoshihisa Shinagawa, Gerardo Hermosillo, Chaowei Wu
Hugging Face Trending Papers
Jul 27

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations.

arXiv AI
Aug 5

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.

By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv AI
Sep 25

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold