arXiv Machine Learning By Zhengyu Wu, Xu Wang, Hongchao Qin, Xunkai Li, Guang Zeng, Rong-Hua Li, Guoren Wang

Multimodal Graph Negative Learning

Read the original on arXiv Machine Learning →

arXiv:2606. 12863v2 Announce Type: replace Abstract: Multimodal attributed graphs (MAGs) integrate graph topology with heterogeneous modality attributes, such as text and images, thereby enabling richer modeling of complex relational systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 22

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

arXiv:2607. 19128v1 Announce Type: new Abstract: Vision-language models (VLMs) provide a unified representation space for textual and visual information, yet their potential as general-purpose backbones for graph-structured data remains largely unexplored.

By Jiayi Yang, Yifang Chen, Yuanfu Sun, Jiajin Liu, Qiaoyu Tan
arXiv Machine Learning
1d ago

MOVE: Multimodal Open-world Verification and Expansion for Graph Learning

The paper introduces MOVE, a framework for multimodal open‑world verification and expansion in graph learning. MOVE jointly uses visual tokens, textual attributes, and graph context to identify nodes that cannot be assigned to existing classes, then employs a multimodal LLM to generate candidate class descriptions. It selectively expands the class space only when multimodal evidence consistently supports the new classes, avoiding redundancy, and reports an average 11.87% improvement across unknown recognition, open‑domain annotation, and downstream graph learning tasks.

By Zekai Chen, Jiayang Xing, Xun Wu, Miao Zhang, Xunkai Li, Kairui Yang, Zhengyu Wu, Xu Wang, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Aug 27

Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?

The paper examines a multimodal approach that combines a self‑supervised GNN encoder with an alternating optimization scheme involving a language‑model teacher. Despite the expectation that this joint strategy would enhance predictive performance, the authors find that the combined model fails to deliver significant gains. They identify six key factors—ranging from anchor strength trade‑offs to misaligned representation spaces—that explain why the integration of text knowledge does not fully benefit graph learning.

By Fumiaki Kimino (SOKENDAI), Ryoma Sato (SOKENDAI, National Institute of Informatics)