arXiv:2608. 00623v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains.
By Yinlin Zhu, Di Wu, Yi Zhang, Xunkai Li, Wang Luo, Wei-Jin Huang, Miao Hu, Guocong Quan
arXiv:2606. 12867v2 Announce Type: replace Abstract: Multimodal-attributed graphs (MAGs) couple graph topology with node semantics from text, images, and other modalities.
By Zhengyu Wu, Xu Wang, Hongchao Qin, Xunkai Li, Guang Zeng, Rong-Hua Li, Guoren Wang
arXiv:2606. 14172v1 Announce Type: new Abstract: Multimodal Attributed Graphs (MAGs) model real-world entities by coupling graph topology with heterogeneous attributes such as text and images.
By Sirui Zhang, Xu Wang, Zhengyu Wu, Xunkai Li, Hongchao Qin
arXiv:2609.06668v1 Announce Type: new
Abstract: Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure an...
By Sirui Zhang, Yubing Zhou, Xunkai Li, Zekai Chen, Shumeng Li, Wang Luo, Yinlin Zhu, Yujin Gao, Rong-Hua Li
arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.
By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
MURAL is a multimodal recommendation framework that replaces static similarity graphs with a dynamic topology discovery process. It uses an Adaptive Edge Learner to find latent item-item correlations efficiently and an Uncertainty-Aware Fusion module to down‑weight noisy modality signals based on aleatoric uncertainty. The model also incorporates a contrastive teacher‑student alignment to stabilize training and has been shown to outperform state‑of‑the‑art baselines on large‑scale TikTok and Amazon datasets, providing both higher accuracy and interpretability.
By Ahmad Mousavi (Department of Mathematics,Statistics American University), Majid Alikhani (Independent Researcher), Yeon-Chang Lee (Department of Computer Science,Engineering Ulsan National Institute of Science,Technology), Roberto Corizzo (Department of Computer Science American University), Yeganeh Abdollahinejad (Department of Biosystems,Agricultural Engineering Michigan State University)
arXiv:2606. 09301v1 Announce Type: new Abstract: Multimodal federated graph learning (MM-FGL) aims to collaboratively learn from decentralized graphs with text and images.
By Zekai Chen, Miao Zhang, Jiayang Xing, Xunkai Li, Xun Wu, Rong-Hua Li, Guoren Wang
arXiv:2608.24795v1 Announce Type: new
Abstract: Recently, the rapid advancement of multimodal domains has driven a data-centric paradigm shift in graph ML, transitioning from text-attributed to multi...
By Xunkai Li, Zekai Chen, Zhengyu Wu, Henan Sun, Daohan Su, Guang Zeng, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv:2607. 26023v1 Announce Type: new Abstract: Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks.
By Ankang Yang, Jitao Zhao, Di Jin, Yuxiao Huang, Dongxiao He
arXiv:2506. 02568v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated substantial efficacy in advancing graph-structured data analysis.
By Dongzhe Fan, Yi Fang, Jiajin Liu, Djellel Difallah, Qiaoyu Tan
arXiv:2607. 23245v1 Announce Type: cross Abstract: Multimodal Federated Learning is often challenged by arbitrary modality missingness and Non-IID data distributions, which lead to severe representation drift and hinder effective collaboration across clients.
By Haochen Liang, Jie Zhang, Hideya Ochiai
GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.
By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang