The paper investigates combining a domain‑specific self‑supervised task—voxel‑level brain age prediction—with a general task—image inpainting—to pretrain models for brain MRI segmentation. A multitask pretraining framework jointly optimizes both objectives, yielding representations that outperform single‑task pretraining and training from scratch on three segmentation benchmarks (multiple sclerosis lesions, ischemic stroke lesions, and cortical structures). The study demonstrates that integrating domain‑specific and general self‑supervised tasks benefits the development of generalizable neuroimaging foundation models.
By Tasneem Nasser, Susanne Schmid, Roberto Souza, Naser El-Sheimy
arXiv:2608. 06122v1 Announce Type: cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series.
By Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
arXiv:2504. 06299v2 Announce Type: replace-cross Abstract: Multimodal prediction models based on imaging and clinical data are increasingly used for clinical decision support, yet their interpretability remains limited.
By Lisa Herzog, Jonas Br\"andli, Maurice Schneeberger, Loran Avci, Nordin Dari, Martin H\"ansel, Hakim Baazaoui, Pascal B\"uhler, Susanne Wegener, Beate Sick
arXiv:2607. 09892v1 Announce Type: cross Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer.
By Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao
arXiv:2607. 17782v1 Announce Type: cross Abstract: Foundation models pretrained using self-supervised learning have transformed computer vision by learning transferable representations from large-scale unlabeled data.
By Moona Mazher, Abdul Qayyum, Steven A. Niederer, Daniel C. Alexander
arXiv:2604. 27277v3 Announce Type: replace-cross Abstract: Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data.
By Yizhou Wu, Shansong Wang, Yuheng Li, Mojtaba Safari, Mingzhe Hu, Chih-Wei Chang, Harini Veeraraghavan, Xiaofeng Yang
The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv:2605. 23995v4 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning transferable representations from unlabeled data.
By Chathura Wimalasiri, Kishor Nandakishor, Marimuthu Palaniswami
arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.
By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
X‑LMC is a spatiotemporal deep‑learning framework that automatically scores leptomeningeal collateral (LMC) status from time‑resolved biplane digital subtraction angiography (DSA). It uses a DINOv2 backbone to encode spatial frames, a token‑level cross‑view attention module to fuse orthogonal projections, and a recurrent network to model contrast bolus dynamics. On a multicenter dataset of 134 M1‑segment occlusion patients, X‑LMC achieved a Quadratic Weighted Kappa of 0.398 and a macro‑F1 of 0.711, outperforming static and other spatiotemporal baselines and matching clinical inter‑rater agreement.
arXiv:2605. 23995v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has emerged as a promising paradigm for addressing the annotation bottleneck in medical imaging by learning representations from unlabeled data.
By Chathura Wimalasiri
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.