arXiv Machine Learning

Context-aware Modality-Topology Co-Alignment for Multimodal Attributed Graphs

arXiv:2606. 14172v1 Announce Type: new Abstract: Multimodal Attributed Graphs (MAGs) model real-world entities by coupling graph topology with heterogeneous attributes such as text and images.

arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Sep 25

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.

By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Aug 4

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

arXiv:2608. 00623v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains.

By Yinlin Zhu, Di Wu, Yi Zhang, Xunkai Li, Wang Luo, Wei-Jin Huang, Miao Hu, Guocong Quan
arXiv AI
Jul 14

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.

By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
arXiv Machine Learning
Sep 7

MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning

MURAL is a multimodal recommendation framework that replaces static similarity graphs with a dynamic topology discovery process. It uses an Adaptive Edge Learner to find latent item-item correlations efficiently and an Uncertainty-Aware Fusion module to down‑weight noisy modality signals based on aleatoric uncertainty. The model also incorporates a contrastive teacher‑student alignment to stabilize training and has been shown to outperform state‑of‑the‑art baselines on large‑scale TikTok and Amazon datasets, providing both higher accuracy and interpretability.

By Ahmad Mousavi (Department of Mathematics,Statistics American University), Majid Alikhani (Independent Researcher), Yeon-Chang Lee (Department of Computer Science,Engineering Ulsan National Institute of Science,Technology), Roberto Corizzo (Department of Computer Science American University), Yeganeh Abdollahinejad (Department of Biosystems,Agricultural Engineering Michigan State University)
arXiv AI
Jun 8

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

arXiv:2602. 07026v3 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions.

By Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, Chonghan Liu, Hanzhen Zhao, Chen Liu, Xiaoxing Hu, Ziyue Qiao, Hao Tang, Xiaobin Hu, Chengwei Qin, Hui Xiong, Yu Qiao, Shuicheng Yan