arXiv AI

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

arXiv AI
Aug 11

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

arXiv:2608. 08135v1 Announce Type: cross Abstract: Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task.

By Daniele Molino, Alessio Zoboli, Camillo Maria Caruso, Valerio Guarrasi, Paolo Soda
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv Machine Learning
Aug 18

Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.

By Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando P\'erez-Garc\'ia
arXiv AI
Jul 28

scMIR: a vision-language foundation model for single-cell light microscopy image representation

arXiv:2607. 22712v1 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis.

By Yifan Shang (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China, College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Jiahui Tan (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Xiangxiang Zeng (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Renjie Zhou (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China)
arXiv Computer Vision
Aug 28

Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation

The paper introduces an unsupervised domain adaptation framework that aligns redundancy-reducing features to enable accurate 3D segmentation of cone-beam CT (CBCT) without target-domain annotations or inference-time adaptation. The method is architecture-agnostic, working with both CNN-based and ViT-based foundation models, and is evaluated on two liver segmentation benchmarks for interventional vascular procedures and radiation therapy. Results show that even large pretrained segmentation networks need explicit feature-space bridging to generalize across diagnostic CT and CBCT, and the proposed approach consistently outperforms existing pretrained foundation models and UDA strategies.

By Gauthier Miralles, Loic Le Folgoc, Vincent Jugnon, Pietro Gori
arXiv Computer Vision
Sep 22

AdaptiveCDM: Source-Free Few-Shot Domain Adaptation for Cell Detection in Microscopic Images

AdaptiveCDM is a modular framework for source‑free few‑shot domain adaptation in cell detection, enabling a pretrained model to adapt to new imaging domains using only a handful of labeled target images and no source data. It combines Resolution‑Aware Augmentation (RAug) to balance scarce, class‑imbalanced samples while preserving cellular morphology, and Category‑Aware Representation Learning (CARL) to strengthen class‑consistent proposals for better localization and classification. Experiments on M5 and Raabin‑WBC datasets show that AdaptiveCDM achieves competitive or superior mAP scores compared to state‑of‑the‑art methods under their respective supervision settings.

By Nimra Dilawar, Sara Nadeem, Javed Iqbal, Waqas Sultani, Mohsen Ali
arXiv Computer Vision
Sep 24

nnFoundation: 3D Foundation Models for Radiology

nnFoundation introduces complementary convolutional and transformer-based 3D foundation models for radiology, trained on 2.1 million CT, MRI, and PET volumes from 125 datasets. The models are evaluated on 108 tasks—including segmentation, detection, classification, report generation, and image retrieval—under domain shift, low-data, and low-compute scenarios, consistently outperforming prior 3D foundation models and training from scratch. Performance varies by task type, with convolutional models excelling at spatially localized tasks and transformer models at global semantic reasoning, and dynamic alignment with dataset characteristics further enhances transferability.

By Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, B\'alint Kov\'acs, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica K\"achele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Kr\"amer, Maximilian Zenk, Tim R\"adsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David K\"ugler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina Decker, Mohammad Yaqub, Arnd Heuser, Martin Reuter, Julia A. Schnabel, Tobias Heimann, Florin Ghesu, Paul Brachmann, Claus P. Heu{\ss}el, Alexander Radbruch, Gianluca Brugnara, Aditya Rastogi, Martha Foltyn-Dumitru, Heinz-Peter Schlemmer, Ignaz Reicht, Julius C. Holzschuh, Michael Bach, Bram Stieltjes, Kai Schlamp, Lena Maier-Hein, Marco Nolden, Ralf Floca, Paul F. J\"ager, Philipp Vollmuth, Fabian Isensee, Klaus H. Maier-Hein
arXiv Computer Vision
Sep 2

Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications

Pix2Rep-v2 is a self‑supervised learning framework that learns pixel‑ and voxel‑level representations for dense medical imaging tasks, using a redundancy‑reduction objective and equivariance principles to scale to 3D and wide field‑of‑view data. The method is evaluated on four datasets across multiple modalities, tasks, and backbones, demonstrating higher data‑efficiency in few‑shot scenarios and competitive performance, such as a +9.3 Dice point improvement in one‑shot segmentation on the M&Ms‑2 dataset. An in‑context dense prototype approach is also proposed, eliminating the need for downstream training.

By S. Sifaoui, E. Angelini, S. Toupin, T. Pezel, L. Le Folgoc