The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.
By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
The paper introduces FedCORE, a federated adaptation framework for multimodal graph foundation models that jointly optimizes perception (Encoder) and reasoning (GNN) modules via a shared low‑dimensional latent state. Unlike prior methods that freeze the Encoder, FedCORE allows both components to adapt together, addressing the dependency between multimodal evidence extraction and graph‑based relational reasoning. Experiments show that FedCORE significantly narrows the Encoder–GNN pairing gap, achieving an 80.7% reduction compared to independent joint adaptation.
By Zekai Chen, Xun Wu, Hailin Zhang, Xunkai Li, Yu Liu, Kairui Yang, Muyan Huang, Xuaner Chen, Rong-Hua Li, Guoren Wang
arXiv:2508. 16159v2 Announce Type: replace-cross Abstract: Meta-learning aims to uniformly sample homogeneous support-query pairs, characterized by the same categories and similar attributes, and extract useful inductive biases through identical network architectures.
By Jiaqi Ma, Guo-Sen Xie, Fang Zhao, Zechao Li
arXiv:2605. 15888v2 Announce Type: replace-cross Abstract: Heterogeneous Graph Prompt Learning (HGPL)has emerged as a promising paradigm for bridging the gap between the objectives of pre-training foundation models and their downstream applications in heterogeneous graph settings.
By Peiyuan Li, Yongqi Huang, Jitao Zhao, Dongxiao He, Di Jin, Weixiong Zhang
CauseCollab is a causal unified and modality‑agnostic network designed to improve collaborative perception across heterogeneous sensor modalities. It disentangles semantic factors from modality‑specific confounders using causal metric learning and employs a context‑guided Unified Converter to maintain cross‑modal semantic consistency. The approach requires only minimal adapter training when adding new modalities and achieves state‑of‑the‑art results on the OPV2V and DAIR‑V2X datasets, especially in scenarios with large modality gaps.
By Weize Li, Yang Li, Quan Yuan, Xiaoyuan Fu, Guiyang Luo, Jinglin Li
arXiv:2606. 25665v1 Announce Type: new Abstract: Domain generalization (DG) aims to learn a model from one or more source domains that generalizes to an unseen target domain without accessing target data during training.
By Tien-Hung Nguyen, Tien-Dat Tran, M. -Duong Nguyen, Kok-Seng Wong
arXiv:2604. 07753v2 Announce Type: replace-cross Abstract: Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts.
By Xiangyue Liu, Zijian Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Ping Tan
arXiv:2606. 05109v1 Announce Type: new Abstract: To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information.
By Vasiliki Rizou, Pascal Frossard, Dorina Thanou
The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to jointly perform heterogeneous classification and regression tasks on the Cityscapes dataset, using an InfoMax principle to handle mixed discrete and continuous semantic variables. It compares the new framework against independent single‑task training, conventional task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom, and studies how the capacity of the common unit affects joint task performance. Extensive evaluations show that CMT‑SemCom outperforms all benchmarks.
By Ahmad Halimi Razlighi, Mohammad Siddiqur Rahman, Maximilian H. V. Tillmann, Edgar Beck, Armin Dekorsy
The paper extends the Cooperative Multi-Task Semantic Communication (CMT‑SemCom) framework to handle heterogeneous classification and regression tasks on the Cityscapes dataset. By incorporating the InfoMax principle, the system accommodates mixed discrete and continuous semantic variables, and it is benchmarked against single‑task training, task‑agnostic digital transmission, and single‑encoder multi‑decoder SemCom. Experiments show that CMT‑SemCom outperforms all baselines and provide insights into how the common unit capacity affects joint task performance.
arXiv:2608. 13911v1 Announce Type: new Abstract: Federated multimodal medical AI faces modality heterogeneity at both the client and sample levels: clients may systematically lack access to specific modality types, while individual records within the same client may contain different partial modality subsets.
By Adiba Orzikulova, Dong Min Kim, Jaehong Yoon, Sung-Ju Lee
arXiv:2609.06668v1 Announce Type: new
Abstract: Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure an...
By Sirui Zhang, Yubing Zhou, Xunkai Li, Zekai Chen, Shumeng Li, Wang Luo, Yinlin Zhu, Yujin Gao, Rong-Hua Li