arXiv:2606. 30291v1 Announce Type: new Abstract: Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks.
By Zhifei Hu, Alexandra I. Cristea
GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.
By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang
Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks. However, existing fusion methods often treat text and structure as separate inputs in a shallow, one-way pipeline, which limits deep interaction between modalities and weakens performance under sparse connectivity or cross-graph generalisation.
arXiv:2512. 12477v2 Announce Type: replace Abstract: Estimating node importance in heterogeneous knowledge graphs is a fundamental problem underlying recommendation, search, and knowledge decision systems.
By Jiawen Chen, Yanyan He, Qi Shao, Mengli Wei, Duxin Chen, Wenwu Yu, Yanlong Zhao
arXiv:2606. 06225v1 Announce Type: cross Abstract: Collaborative filtering and graph-based recommendation models are highly effective because they leverage observed user interactions, but this dependence creates a fundamental cold-start challenge when newly added content has no interaction history.
By Anh Truong, John Trenkle, Yuanbo Chen, Honghong Zhao, Abdullah Alchihabi, Effy Fang, Michael Tamir
PEARL is a new framework for inductive knowledge graph completion that treats relational paths as context-conditioned reasoning signals. It builds a query‑specific contextual subgraph from the query entities’ neighborhoods and uses a large language model‑guided retriever to select semantically relevant paths. By constructing a bipartite interaction graph over paths, contextual entities, and a global subgraph representation, and applying a dual‑view contrastive objective, PEARL adapts path embeddings to local and global structural evidence, achieving the best average Hits@10 on WN18RR, FB15k‑237, and NELL‑995.
By Yunchi Yang, Longlong Li, Cunquan Qu
HMGCLIP is a unified multimodal embedding framework that uses a heterogeneous hypergraph to capture both fine‑grained and coarse‑grained product attributes. By mining structure‑aware hard negatives and aligning multi‑granular semantics at relation and hyperedge levels, it enables a dual‑granularity inference mechanism that dynamically fuses attribute evidence. Experiments on a new fine‑grained e‑commerce dataset and the public MAVE benchmark show that HMGCLIP outperforms strong multimodal encoders, MLLMs, and e‑commerce baselines.
By Qiuyu Zhu, Yi Gao, Zhichao Wan, Mingyang Ma
arXiv:2609.21164v1 Announce Type: new
Abstract: Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decisio...
By Inyoung Choi, Sukwon Yun, Jiayi Xin, Jie Peng, Tianlong Chen, Qi Long
MURAL is a multimodal recommendation framework that replaces static similarity graphs with a dynamic topology discovery process. It uses an Adaptive Edge Learner to find latent item-item correlations efficiently and an Uncertainty-Aware Fusion module to down‑weight noisy modality signals based on aleatoric uncertainty. The model also incorporates a contrastive teacher‑student alignment to stabilize training and has been shown to outperform state‑of‑the‑art baselines on large‑scale TikTok and Amazon datasets, providing both higher accuracy and interpretability.
By Ahmad Mousavi (Department of Mathematics,Statistics American University), Majid Alikhani (Independent Researcher), Yeon-Chang Lee (Department of Computer Science,Engineering Ulsan National Institute of Science,Technology), Roberto Corizzo (Department of Computer Science American University), Yeganeh Abdollahinejad (Department of Biosystems,Agricultural Engineering Michigan State University)
arXiv:2606. 11898v1 Announce Type: cross Abstract: Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenarios, such as citation networks, e-commerce platforms, social media, and web pages.
By Hengyi Feng, Zeang Sheng, Meiyi Qiang, Meiyi Qiang, Wentao Zhang
arXiv:2606. 12245v1 Announce Type: cross Abstract: Cold-start item recommendation remains a persistent challenge in real-world systems due to the absence of interaction histories.
By Kangning Zhang, Yingjie Qin, Weinan Zhang, Yong Yu, Jianghao Lin
Although recent Multimodal Large Language Models (MLLMs) have advanced general product understanding, they implicitly encode product information into global embeddings, thereby limiting their ability...