arXiv:2606. 01856v1 Announce Type: cross Abstract: Multimodal Federated Learning (MMFL) enables privacy-preserving collaborative learning across decentralized clients with heterogeneous data and modality availability.
By Zixin Zhang, Fan Qi, Shuai Li, Xiaoshan Yang, Changsheng Xu
arXiv:2512. 02076v2 Announce Type: replace-cross Abstract: We propose FDRMFL, a task-driven multimodal feature extraction framework for federated regression under non-IID data distributions.
By Haozhe Wu
arXiv:2607. 27289v1 Announce Type: new Abstract: The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction.
By Yu Chang, Anzhe Cheng, Chenwei Wu, Zhuoran Wang, Jiahao Chen, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Paul M. Thompson, Liyue Shen, Paul Bogdan
arXiv:2607. 23245v1 Announce Type: cross Abstract: Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients.
By Haochen Liang, Jie Zhang, Hideya Ochiai
arXiv:2607. 06633v1 Announce Type: cross Abstract: In this paper, we address the problem of multimodal federated learning with missing modality.
By Aavash Chhetri, Bibek Niroula, Eduard Vazquez, Yash Raj Shrestha, Prashnna Gyawali, Loris Bazzani, Binod Bhattarai
FedCoRe is a federated learning framework that addresses missing modalities in healthcare by learning representation- or logit-space corrections instead of generating synthetic data. In a MIMIC-derived respiratory deterioration task, the method uses paired examples where a modality is present during training but may be absent at deployment, allowing only those clients to update the completion module. Experiments show that FedCoRe can recover roughly half of the performance lost when ECG or CXR data are hidden, but the framework should only be deployed when paired examples and validation evidence confirm the modality’s presence.
arXiv:2608. 15310v1 Announce Type: cross Abstract: Multimodal data collected by heterogeneous devices are used for collaborative training, where federated learning (FL) serves as a key paradigm for effective distributed modeling with data privacy preservation.
By Zhenyan Liu, Hua Zhang, Haoran Gao, Qi Li, Hongliang Zhu, Huiyu Zhou, Zongliang Shen, Yanxin Xu, Jiahui Wang
arXiv:2608. 13911v1 Announce Type: new Abstract: Federated multimodal medical AI faces modality heterogeneity at both the client and sample levels: clients may systematically lack access to specific modality types, while individual records within the same client may contain different partial modality subsets.
By Adiba Orzikulova, Dong Min Kim, Jaehong Yoon, Sung-Ju Lee
arXiv:2607. 20742v1 Announce Type: new Abstract: Multimodal learning is a robust approach to improve predictive performance in applications such as medical prognosis.
By Mohammad Raahemi, Ali Sekhavati, Alireza Maleki, Hamid Nasiri
FedCoRe is a federated learning framework that addresses missing modalities in healthcare by learning representation- or logit-space corrections instead of generating synthetic data. In a MIMIC-derived respiratory deterioration task, the method uses paired examples where a modality is present or absent to train a completion module, achieving partial recovery of performance lost when modalities like ECG or CXR are hidden. The approach emphasizes validation-gated deployment, ensuring that completion is only applied when paired examples and validation evidence support the presence of the missing modality.
By Holger R. Roth, Ziyue Xu, Peter Cnudde
arXiv:2606. 15743v1 Announce Type: new Abstract: This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature embeddings in a task-independent manner before supervised prediction.
By Hassan Ismkhan, Hamid Bouchahcia
The paper identifies that in multimodal learning, optimization often produces asymmetric certainty gains, with the stronger modality becoming more confident than the weaker one, which leads to imbalanced contributions and suboptimal performance. The authors attribute this issue to unimodal characteristics and propose a Max Confidence Regularization (MaxCR) method that tracks each modality’s semantic confidence via a nonlinear sparsity measure and applies max suppression and excitation to balance confidence levels. Experiments on standard datasets demonstrate that MaxCR improves overall performance compared to state‑of‑the‑art multimodal baselines.
By Longfei Huang, Xiangyu Wu, Yang Yang