arXiv AI By Zhemin Zhang, Weijie Chen, David Le, Amara Tariq, Alex Wallace, Matthew Stib, Juan Maria Farina, Chadi Ayoub, Reza Arsanjani, Imon Banerjee

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

Read the original on arXiv AI →

arXiv:2606. 15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 27

Autoregressive EHR Foundation Models with Multimodal Inputs

arXiv:2607. 22264v1 Announce Type: new Abstract: Autoregressive foundation models trained on tokenized electronic health records (EHRs) can support zero-shot clinical prediction, yet most operate on structured event codes alone, and do not incorporate multiple modalities in a principled way.

By Yuxuan Liu, Joshua Placidi, Jinpei Han, Alfred John Balston, Marek Rei, A. Aldo Faisal
arXiv AI
Jun 17

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.

By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
arXiv Computer Vision
Sep 7

Real-World Multi-Modal and Longitudinal Lung Cancer Dataset

The paper presents a newly curated, multi-center, multi-modal, and longitudinal lung cancer dataset comprising 1,365 patients with whole-slide images, CT scans, PET scans, structured clinical data, transcriptomics, and follow-up information. The dataset features substantial, non-uniform missingness across modalities, making it ideal for evaluating robust multi-modal fusion strategies. Benchmarks on 12‑month overall survival, disease‑specific survival, and longitudinal hazard prediction demonstrate that integrating complementary modalities consistently outperforms uni-modal approaches, even under severe missing data.

By Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago, Catarina Barata
arXiv AI
Sep 21

MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention

MIST is a multimodal survival prediction framework that fuses whole-slide images and genomic profiles by representing genomic features as tokens that query histology context tokens derived from a foundation model. The architecture enriches molecular information with histology context before survival prediction, avoiding late-stage merging of separately encoded modalities. Training incorporates discrete-time survival prediction, genomic feature masking, WSI dropout, and contrastive alignment, and demonstrates improved external C-index across colon, renal, lung, and glioblastoma cohorts compared to standard fusion baselines.

By Muhammet Sami Yavuz, Sabri Mustafa Kahya, Richard R. Chen, Jana Lipkova, Benedikt Wiestler