MMAP is a Multimodal Missing‑Aware Alignment Pretraining method designed to learn image‑tabular representations from incomplete data. It uses a sigmoid contrastive learning image encoder with generative reconstruction, a tabular encoder based on a foundation model, and a missing token generator to handle missing modalities. The approach is evaluated on longitudinal Alzheimer’s tasks—predicting disease stage conversion and amyloid status—and outperforms both multimodal and unimodal baselines.
By Fiona Kekwick, Matthew Baugh, Bernhard Kainz, Paul M. Matthews, Wenjia Bai
arXiv:2602. 18400v3 Announce Type: replace-cross Abstract: Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice.
By Junkai Liu, Nay Aung, Theodoros N. Arvanitis, Joao A. C. Lima, Steffen E. Petersen, Le Zhang
arXiv:2606. 17989v1 Announce Type: cross Abstract: Multi-contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis.
By Yonghao Chen, Sicheng Yang, Rui Tang, Lei Zhu
arXiv:2606. 15743v1 Announce Type: new Abstract: This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature embeddings in a task-independent manner before supervised prediction.
By Hassan Ismkhan, Hamid Bouchahcia
arXiv:2606. 06328v1 Announce Type: new Abstract: In healthcare, multimodal time series tasks often operate on incomplete observations in practice, for example when ECG segments are lost because electrodes detach or an entire respiratory channel is unavailable during overnight monitoring.
By Ziwen Kan, Wugeng Zheng, Tianlong Chen, Song Wang
arXiv:2606. 12362v1 Announce Type: cross Abstract: We study multimodal learning under missing modalities, with particular motivation from bioscience applications in which heterogeneous modalities are often only partially available when decisions need to be made.
By Hui Wang, Tianyu Ren, Joseph Butler, Christopher Baker, Karen Rafferty, Simon McDade
arXiv:2606. 09907v1 Announce Type: cross Abstract: Multimodal clinical learning is increasingly important for integrating diverse patient data, including imaging, text, and personalised health records.
By Maxx Richard Rahman, Prakhar Kumar, Wolfgang Maass
arXiv:2607. 14995v1 Announce Type: new Abstract: Multimodal Contrastive Learning (CL) has shown significant performance in aligning representations across various data modalities and improving downstream tasks, especially in healthcare.
By Sara Ketabi, Matthias W. Wagner, Cynthia Hawkins, Uri Tabori, Birgit Betina Ertl-Wagner, Farzad Khalvati
The paper introduces Neuro‑JEPA, a sparse multimodal foundation model that learns unified representations of brain MRI across T1w, T2w, and FLAIR sequences using a latent predictive objective and a Mixture‑of‑Experts architecture. It was pretrained on over 1.5 million scans from 428,647 studies and systematically evaluates architectural, masking, objective, and sparsity choices for robust multimodal representation learning. Across 47 tasks from three health systems and 12 public datasets, Neuro‑JEPA consistently outperforms a simple CNN baseline, demonstrating its effectiveness for both clinical and research applications.
By Haoxu Huang, Long Chen, Jingyun Chen, Jinu Hyun, James Ryan Loftus, Kara Melmed, Daniel Orringer, Jennifer Frontera, Seena Dehkharghani, Arjun Masurkar, Narges Razavian
arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.
By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.
arXiv:2607. 09892v1 Announce Type: cross Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer.
By Chicago Y. Park, Jialin Mao, Xiaojian Xu, Taha Kass-Hout, Ulugbek S. Kamilov, Cao Xiao