arXiv Computer Vision

Real-World Multi-Modal and Longitudinal Lung Cancer Dataset

The paper presents a newly curated, multi-center, multi-modal, and longitudinal lung cancer dataset comprising 1,365 patients with whole-slide images, CT scans, PET scans, structured clinical data, transcriptomics, and follow-up information. The dataset features substantial, non-uniform missingness across modalities, making it ideal for evaluating robust multi-modal fusion strategies. Benchmarks on 12‑month overall survival, disease‑specific survival, and longitudinal hazard prediction demonstrate that integrating complementary modalities consistently outperforms uni-modal approaches, even under severe missing data.

arXiv AI
Jun 17

Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis

arXiv:2606. 17115v1 Announce Type: cross Abstract: Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored.

By Jingyu Hu, Giuseppe Tripodi, Reed Naidoo, Sarah F. McGough, Tapabrata Chakraborti
arXiv Machine Learning
Jul 24

Multimodality Stacking with Blockwise missing values and application to the PIONeeR biomarkers study for prediction of resistance to immunotherapy

arXiv:2605. 25050v2 Announce Type: replace-cross Abstract: Integrating multimodal datasets in clinical oncology is frequently hindered by high dimensionality and blockwise missingness, where entire data sources are unavailable for specific patient subsets.

By Mohamed Boussena, Florence Monville, Jacques Fieschi-Meric, Frederic Vely, Pierre Milpied, Julien Mazieres, Maurice Perol, Eric Vivier, Laurent Greillier, Fabrice Barlesi, Sebastien Benzekry
Hugging Face Trending Papers
Jul 9

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the scarcity of curated imaging cohorts with reliable outcome data. This study evaluates whether representations from a domain-specific foundation model can be used for multimodal survival prediction in data-constrained clinical settings.

arXiv Machine Learning
Aug 26

A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology

arXiv:2608.24688v1 Announce Type: new Abstract: Precision oncology necessitates a longitudinal model of patient state that captures cancer evolution and treatment over time, integrating multimodal ob...

By Eugene Vorontsov, Yi Kan Wang, Alican Bozkurt, Adam Casson, Ludmila Tydlitatova, Michal Zelechowski, Ezra E. W. Cohen, Jyoti D. Patel, Max Banaszak, Caitlin McWilliams, Shane Colley, Kate Sasser, Ryan Fukushima, Eric Lefkofsky, Razik Yousfi, Siqi Liu
arXiv Machine Learning
Aug 19

MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology

MultiSigBERT is a unified framework that performs multimodal sequential survival modeling in oncology by integrating narrative clinical reports, numerical measurements, and structured variables. The method converts free-text reports into sentence embeddings, compresses them with modality-specific PCA, and concatenates them with structured covariates to create joint temporal trajectories. These trajectories are encoded using the Signature transform from Rough Paths theory, and the resulting high-dimensional features are fed into a LASSO-regularized Cox model, achieving a concordance index of 0.743 on an independent test set of over 2,500 patients.

By Paul Minchella, St\'ephane Chr\'etien, Guillaume Metzler, Lo\"ic Verlingue, R\'emi Vaucher