arXiv:2606. 10255v1 Announce Type: cross Abstract: Cryo-electron tomography (cryoET) has emerged as a powerful tool in structural and cellular biology by enabling direct visualization of macromolecular structures within intact cells, thereby linking molecular architecture to cellular organization in a native context.
By Jonathan Schwartz, Utz Heinrich Ermel, C. Braxton Owens, Zhuowen Zhao, Ariana Peck, Gus L. W. Hart, Grant J. Jensen, Bridget Carragher, Dari Kimanius
The paper introduces CARNIVAL, a model for protein annotation in cryo-electron tomography (cryo-ET) volumes that leverages simulated data and a forward model to generate domain‑specific augmented paired views for self‑supervised training. By incorporating simulation‑derived protein positions and identities into the architecture and loss function, the model localises semantic information at protein locations. CARNIVAL is evaluated on real tomograms without finetuning and outperforms a state‑of‑the‑art contrastive model that lacks forward‑model paired views or privileged information.
By Bogdan Toader, Kiarash Jamali, Tanmay A. M. Bharat, Sjors H. W. Scheres
arXiv:2609.37605v1 Announce Type: cross
Abstract: Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-project...
By AmirEhsan Khorashadizadeh, Benjam\'in B\'ejar
arXiv:2609.23019v1 Announce Type: cross
Abstract: Soma instance segmentation, i.e., identifying and delineating individual cell somas as distinct instances, is crucial for cellular analysis and conne...
By Mohammad Khateri, Morteza Ghahremani, Jussi Tohka, Alejandra Sierra
arXiv:2606. 00955v1 Announce Type: new Abstract: Despite the growing availability of cryo-electron microscopy (cryo-EM) density maps, effectively leveraging them for protein representation remains challenging.
By Dan Luo, Xuan Lin, Peng Zhou, Junwen Zhu, Tengfei Ma, Xiangxiang Zeng, Yiping Liu
arXiv:2609.09394v1 Announce Type: new
Abstract: Recovering metric 3D geometry from monocular images is a fundamental computer vision task, yet current methods remain heavily fragmented by fixed camer...
By Botao Ye, Marc Pollefeys, Ming-Hsuan Yang, Abhijit Kundu
arXiv:2608. 04515v1 Announce Type: cross Abstract: Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices.
By Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang
arXiv:2604.19609v2 Announce Type: replace
Abstract: Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong dom...
By Kadir Yilmaz, Adrian Kruse, Tristan H\"ofer, Daan de Geus, Bastian Leibe
MTMed3D is a multi-task Transformer-based model that jointly performs 3D detection, segmentation, and classification in medical imaging. It uses a shared Transformer encoder to produce multi-scale features, with separate CNN decoders for each task. Evaluated on BraTS 2018 and 2019, it achieves strong results, especially in detection, while reducing computational cost and inference time compared to single-task models.
By Fan Li, Arun Iyengar, Lanyu Xu
arXiv:2603. 11625v2 Announce Type: replace-cross Abstract: While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, their deployment for 3D volumetric data remains constrained by significant computational inefficiencies.
By Shengyuan Liu, Zanting Ye, Yunrui Lin, Chen Hu, Wanting Geng, Xu Han, Bulat Ibragimov, Yefeng Zheng, Yixuan Yuan
arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.
By Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando P\'erez-Garc\'ia
arXiv:2608. 08713v1 Announce Type: cross Abstract: Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges.
By Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein