arXiv AI

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

arXiv:2606. 15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.

arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv AI
Jun 17

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.

By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
arXiv Computer Vision
Sep 7

Real-World Multi-Modal and Longitudinal Lung Cancer Dataset

The paper presents a newly curated, multi-center, multi-modal, and longitudinal lung cancer dataset comprising 1,365 patients with whole-slide images, CT scans, PET scans, structured clinical data, transcriptomics, and follow-up information. The dataset features substantial, non-uniform missingness across modalities, making it ideal for evaluating robust multi-modal fusion strategies. Benchmarks on 12‑month overall survival, disease‑specific survival, and longitudinal hazard prediction demonstrate that integrating complementary modalities consistently outperforms uni-modal approaches, even under severe missing data.

By Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago, Catarina Barata
arXiv AI
Sep 21

MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention

MIST is a multimodal survival prediction framework that fuses whole-slide images and genomic profiles by representing genomic features as tokens that query histology context tokens derived from a foundation model. The architecture enriches molecular information with histology context before survival prediction, avoiding late-stage merging of separately encoded modalities. Training incorporates discrete-time survival prediction, genomic feature masking, WSI dropout, and contrastive alignment, and demonstrates improved external C-index across colon, renal, lung, and glioblastoma cohorts compared to standard fusion baselines.

By Muhammet Sami Yavuz, Sabri Mustafa Kahya, Richard R. Chen, Jana Lipkova, Benedikt Wiestler
arXiv Computer Vision
Aug 27

PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology

PANDA (Prototype‑Anchored Data Alignment) is a two‑stage framework that enables a primary‑modality model to benefit from auxiliary modalities even when those modalities are only partially paired or absent at inference. In Stage 1, a shared embedding is learned from the paired subset and class prototypes are estimated from the auxiliary data; in Stage 2, the primary encoder is trained on all subjects using cross‑entropy and alignment to the frozen prototypes. PANDA was evaluated on Alzheimer’s MRI and TCGA‑Lung pathology, achieving significant AUC gains and improved survival prediction while requiring no auxiliary inputs during deployment.

By Sheethal Bhat, Mahfuzur Rahman Chowdhury, Paula Andrea Perez-Toro, Stephan Wunderlich, Rose Dawn Bharat, Siming Bayer, Andreas Maier
arXiv AI
Jun 2

UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment

arXiv:2606. 00170v1 Announce Type: cross Abstract: In recent years, emotion recognition based on physiological signals such as electroencephalogram (EEG) has gained considerable attention, as internal physiological data offer greater objectivity and reliability compared to external behavioral data like facial expressions.

By Zheng Wang, Shuo Wang, Junhong Wang
arXiv AI
Jul 14

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.

By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
arXiv Machine Learning
Jul 20

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

arXiv:2607. 15447v1 Announce Type: new Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods.

By Jingteng Li, Alexander Capstick, Louise Rigny, Iona Biggart, Neil J Sebire, Payam Barnaghi