CoMLP introduces a cooperatively-gated MLP module that fuses multimodal medical data—such as imaging modalities and clinical reports—without relying on computationally heavy cross-attention. The module uses regional and dilated MLP interactions to capture both local and global cross-modal dependencies, enabling fine-grained fusion at high spatial resolutions. Experiments on five segmentation benchmarks, covering 2D/3D images and diverse anatomical regions, show consistent improvements over state-of-the-art multi-modal and language-guided methods, highlighting the effectiveness of MLP-based interaction for medical image segmentation.
By Mingyuan Meng, Shuchang Ye, Mingjian Li, Zhenyu Zhao, Jinman Kim, Lei Bi
arXiv:2606. 15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.
By Zhemin Zhang, Weijie Chen, David Le, Amara Tariq, Alex Wallace, Matthew Stib, Juan Maria Farina, Chadi Ayoub, Reza Arsanjani, Imon Banerjee
arXiv:2608. 02769v1 Announce Type: cross Abstract: Multimodal supervised learning seeks to leverage multiple heterogeneous data sources to improve predictive performance.
By Sagnik Nandy, Samriddha Lahiry, Pragya Sur, Subhabrata Sen
Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity.
Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing surveys mainly summarize task-specific algorithms, datasets, or preprocessing techniques independently, lacking a unified perspective on their co-evolution with modern artificial intelligence.
The paper introduces Fusion Anything Model (FAM), a foundation model designed for generalized multimodal data fusion that can handle arbitrary modality combinations and prediction tasks. FAM is trained on large-scale synthetic multimodal datasets generated via Structural Multimodal Causal Models (SMCMs), enabling it to encode transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show that FAM performs competitively with specialized models without requiring task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.
By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
The paper introduces a generalized multimodal foundation model that can handle arbitrary combinations of modalities and prediction tasks. It trains on large-scale synthetic multimodal datasets with diverse causal structures to learn transferable multimodal correlations. Experiments on 18 real-world datasets across 12 modalities and 11 tasks show competitive performance compared to specialized models without task-specific adaptation.
By Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, Changqing Zhang
arXiv:2607. 24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment.
By Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
arXiv:2606. 02659v1 Announce Type: cross Abstract: Multimodal data fusion involves integrating and analyzing information from multiple modalities to uncover latent correlations and complementary patterns, thereby enhancing data processing and decision-making.
By Dong Li, Lingling Zhang, Binghao Han, Linlin Ding, Yue Kou
arXiv:2607. 11839v1 Announce Type: cross Abstract: This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments.
By Divya Mereddy, Jeevan Beedareddy
arXiv:2603.02767v4 Announce Type: replace-cross
Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield repre...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Yaqian Li, Kun He