arXiv Machine Learning

The microscope is the mask: privileged views and labels from a cryo-ET forward model

The paper introduces CARNIVAL, a model for protein annotation in cryo-electron tomography (cryo-ET) volumes that leverages simulated data and a forward model to generate domain‑specific augmented paired views for self‑supervised training. By incorporating simulation‑derived protein positions and identities into the architecture and loss function, the model localises semantic information at protein locations. CARNIVAL is evaluated on real tomograms without finetuning and outperforms a state‑of‑the‑art contrastive model that lacks forward‑model paired views or privileged information.

arXiv Machine Learning
Jun 10

POPSICLE: Benchmark Datasets for Segmentation and Localization in CryoET

arXiv:2606. 10255v1 Announce Type: cross Abstract: Cryo-electron tomography (cryoET) has emerged as a powerful tool in structural and cellular biology by enabling direct visualization of macromolecular structures within intact cells, thereby linking molecular architecture to cellular organization in a native context.

By Jonathan Schwartz, Utz Heinrich Ermel, C. Braxton Owens, Zhuowen Zhao, Ariana Peck, Gus L. W. Hart, Grant J. Jensen, Bridget Carragher, Dari Kimanius
arXiv AI
4d ago

Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks

Atelier is a self‑supervised framework that uses a transformer‑based hypernetwork to generate implicit neural representations (INRs) for cryo‑EM maps, enabling efficient, scale‑agnostic, coordinate‑conditioned feature extraction. Trained on 5,439 maps from the Electron Microscopy Data Bank, the pretrained INR provides continuous local feature fields that can be used as auxiliary channels for a 3D nested U‑Net, improving voxel‑level property prediction across eight tasks compared to a volume‑only baseline. The approach demonstrates that amortized INRs can serve as a geometry‑aware primitive for large‑scale cryo‑EM analysis.

By Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
arXiv Computer Vision
Sep 4

Using Deep Learning Models Pretrained by Self-Supervised Learning for Protein Localization

The study evaluates self‑supervised learning (SSL) models pretrained on ImageNet‑1k and the Human Protein Atlas (HPA) Field‑of‑View (FOV) for protein localization in microscopy images. DINO‑based Vision Transformer backbones pretrained on either dataset transfer well to the OpenCell dataset, achieving strong performance even without fine‑tuning and improving further when fine‑tuned (0.704 ± 0.027 macro F1 on 17 classes). At the single‑cell level, the HPA‑pretrained model outperforms others in k‑nearest‑neighbor classification across all neighborhood sizes (macro F1 ≥ 0.515).

By Ben Isselmann, Dilara G\"oksu, Heinz Neumann, Andreas Weinmann
arXiv Computer Vision
Sep 2

Pix2Rep-v2: Data-Efficient Representation Learning for Dense Medical Imaging Applications

Pix2Rep-v2 is a self‑supervised learning framework that learns pixel‑ and voxel‑level representations for dense medical imaging tasks, using a redundancy‑reduction objective and equivariance principles to scale to 3D and wide field‑of‑view data. The method is evaluated on four datasets across multiple modalities, tasks, and backbones, demonstrating higher data‑efficiency in few‑shot scenarios and competitive performance, such as a +9.3 Dice point improvement in one‑shot segmentation on the M&Ms‑2 dataset. An in‑context dense prototype approach is also proposed, eliminating the need for downstream training.

By S. Sifaoui, E. Angelini, S. Toupin, T. Pezel, L. Le Folgoc
arXiv Computer Vision
Aug 28

DALE-CT: Depth-Aware 2D Slice Encoders Learn an Anatomical World Model of Chest CT

DALE-CT introduces depth‑aware 2D slice encoders that learn an anatomical world model of chest CT scans without 3D or positional supervision. By sampling self‑supervised views across a physical $z$‑axis slab, the encoder captures how anatomy changes between neighboring slices, enabling it to recover slice ordering and distinguish slices by anatomy alone. The model, trained on a large 287k‑scan corpus, achieves state‑of‑the‑art performance on CT‑RATE and is released with full code and evaluation tools.

By Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner
arXiv AI
Sep 25

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold