arXiv Machine Learning

3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

arXiv:2606. 23964v1 Announce Type: new Abstract: Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells.

arXiv AI
Sep 25

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
arXiv AI
6d ago

Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks

Atelier is a self‑supervised framework that uses a transformer‑based hypernetwork to generate implicit neural representations (INRs) for cryo‑EM maps, enabling efficient, scale‑agnostic, coordinate‑conditioned feature extraction. Trained on 5,439 maps from the Electron Microscopy Data Bank, the pretrained INR provides continuous local feature fields that can be used as auxiliary channels for a 3D nested U‑Net, improving voxel‑level property prediction across eight tasks compared to a volume‑only baseline. The approach demonstrates that amortized INRs can serve as a geometry‑aware primitive for large‑scale cryo‑EM analysis.

By Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
arXiv Computer Vision
Sep 4

Using Deep Learning Models Pretrained by Self-Supervised Learning for Protein Localization

The study evaluates self‑supervised learning (SSL) models pretrained on ImageNet‑1k and the Human Protein Atlas (HPA) Field‑of‑View (FOV) for protein localization in microscopy images. DINO‑based Vision Transformer backbones pretrained on either dataset transfer well to the OpenCell dataset, achieving strong performance even without fine‑tuning and improving further when fine‑tuned (0.704 ± 0.027 macro F1 on 17 classes). At the single‑cell level, the HPA‑pretrained model outperforms others in k‑nearest‑neighbor classification across all neighborhood sizes (macro F1 ≥ 0.515).

By Ben Isselmann, Dilara G\"oksu, Heinz Neumann, Andreas Weinmann
arXiv AI
Jul 28

scMIR: a vision-language foundation model for single-cell light microscopy image representation

arXiv:2607. 22712v1 Announce Type: cross Abstract: Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis.

By Yifan Shang (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China, College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Jiahui Tan (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Xiangxiang Zeng (College of Computer Science and Electronic Engineering, Hunan University, Changsha, China), Renjie Zhou (Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong, China)
arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv Machine Learning
Aug 11

Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis

arXiv:2608. 07575v1 Announce Type: cross Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic: along the under-sampled axial direction the structure can appear discontinuous, hampering reconstruction and automated quantitative analysis.

By Arash Fatehi, Robin Ebbestad, Linus Butt, Hans Blom, Sigrid Lundberg, Hannes Olauson, Hjalmar Brismar, David Unnersj\"o-Jess, Thomas Benzing, Katarzyna Bozek
arXiv Machine Learning
Sep 7

The microscope is the mask: privileged views and labels from a cryo-ET forward model

The paper introduces CARNIVAL, a model for protein annotation in cryo-electron tomography (cryo-ET) volumes that leverages simulated data and a forward model to generate domain‑specific augmented paired views for self‑supervised training. By incorporating simulation‑derived protein positions and identities into the architecture and loss function, the model localises semantic information at protein locations. CARNIVAL is evaluated on real tomograms without finetuning and outperforms a state‑of‑the‑art contrastive model that lacks forward‑model paired views or privileged information.

By Bogdan Toader, Kiarash Jamali, Tanmay A. M. Bharat, Sjors H. W. Scheres