The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.
By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
arXiv:2607. 22712v1 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis.
By Yifan Shang (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China, College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Jiahui Tan (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Xiangxiang Zeng (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Renjie Zhou (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China)
The paper presents MoCoP v2, an enhanced contrastive pretraining method that aligns small molecule embeddings with deep‑learning‑derived cell morphology profiles. By replacing CellProfiler fingerprints with richer image‑encoded features, the new embeddings better capture how molecules alter cell morphology, leading to improved QSAR, toxicity, ADME, and activity predictions. Performance scales log‑linearly with training data size, indicating further gains with larger datasets.
By Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang
arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.
By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
arXiv:2608.07632v2 Announce Type: replace-cross
Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Pa...
By Al\'an F. Mu\~noz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh
arXiv:2608. 16810v1 Announce Type: cross Abstract: Identifying and representing object instances such as cells or nuclei is a common task in microscopy image analysis.
By Ziwen Liu, Martin Weigert
arXiv:2512. 21414v2 Announce Type: replace-cross Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tools.
By Christina Liu, Alan Q. Wang, Joy Hsu, Jiajun Wu, Ehsan Adeli
The study evaluates self‑supervised learning (SSL) models pretrained on ImageNet‑1k and the Human Protein Atlas (HPA) Field‑of‑View (FOV) for protein localization in microscopy images. DINO‑based Vision Transformer backbones pretrained on either dataset transfer well to the OpenCell dataset, achieving strong performance even without fine‑tuning and improving further when fine‑tuned (0.704 ± 0.027 macro F1 on 17 classes). At the single‑cell level, the HPA‑pretrained model outperforms others in k‑nearest‑neighbor classification across all neighborhood sizes (macro F1 ≥ 0.515).
By Ben Isselmann, Dilara G\"oksu, Heinz Neumann, Andreas Weinmann
The paper introduces a self‑supervised, physics‑aware deep learning method for hyperspectral image restoration and super‑resolution in biomedical imaging. It achieves 16× pixel super‑resolution and 12× faster imaging without external training data, preserving biological integrity across synthetic and experimental samples. The approach also reveals disease‑associated metabolic changes and offers physical insights into the model’s workings, with all code released as open source.
By Yuchen Xiang, Zhaolu Liu, Monica Emili Garcia-Segura, Daniel Simon, Boxuan Cao, Vincen Wu, Kenneth Robinson, Yu Wang, Ronan Battle, Najah Sobhan, Robert T. Murray, Xavier Altafaj, John Marshall, Luca Peruzzotti-Jametti, Zoltan Takats
arXiv:2609.09863v1 Announce Type: new
Abstract: Choosing a deep learning architecture for label-free single-cell classification remains an open question, with microscopy benchmarks reporting conflict...
By Philip Graemer, Giuseppe Di Caprio
arXiv:2403. 18026v3 Announce Type: replace-cross Abstract: High-throughput imaging is often constrained by a trade-off between acquisition speed and image quality.
By Dominik Panek, Carina Rz\k{a}ca, Maksymilian Szczypior, Joanna Sorysz, Krzysztof Misztal, Zbigniew Baster, Zenon Rajfur
SynerMedGen is a unified framework that aligns medical multimodal understanding with generation tasks through task alignment. It introduces three generation‑aligned understanding tasks and a two‑stage training strategy that transfers representations learned during understanding to medical image synthesis. The model achieves strong zero‑shot performance on 22 synthesis tasks and outperforms state‑of‑the‑art specialized and unified models when combined with generation training, supported by a new 1M‑sample SynerMed dataset.
By Weiren Zhao, Yi Dong, Cheng Chen