arXiv Machine Learning

Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach

arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Aug 4

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

arXiv:2608. 00623v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains.

By Yinlin Zhu, Di Wu, Yi Zhang, Xunkai Li, Wang Luo, Wei-Jin Huang, Miao Hu, Guocong Quan
arXiv Machine Learning
Jul 22

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

arXiv:2607. 19128v1 Announce Type: new Abstract: Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored.

By Jiayi Yang, Yifang Chen, Yuanfu Sun, Jiajin Liu, Qiaoyu Tan
arXiv Machine Learning
Sep 7

MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning

MURAL is a multimodal recommendation framework that replaces static similarity graphs with a dynamic topology discovery process. It uses an Adaptive Edge Learner to find latent item-item correlations efficiently and an Uncertainty-Aware Fusion module to down‑weight noisy modality signals based on aleatoric uncertainty. The model also incorporates a contrastive teacher‑student alignment to stabilize training and has been shown to outperform state‑of‑the‑art baselines on large‑scale TikTok and Amazon datasets, providing both higher accuracy and interpretability.

By Ahmad Mousavi (Department of Mathematics,Statistics American University), Majid Alikhani (Independent Researcher), Yeon-Chang Lee (Department of Computer Science,Engineering Ulsan National Institute of Science,Technology), Roberto Corizzo (Department of Computer Science American University), Yeganeh Abdollahinejad (Department of Biosystems,Agricultural Engineering Michigan State University)