arXiv Computer Vision

Hub-Spectral Activation of Latent Multimodal Knowledge

Hugging Face Trending Papers
Jun 3

COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations

Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.

arXiv Machine Learning
Sep 25

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.

By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Jul 20

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.

By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
arXiv Machine Learning
Sep 25

ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

ICE (Interaction-aware Clifford Encoder) is a multimodal graph foundation model that uses a node-indexed Clifford latent field to encode topology, text, and images into explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods, preserving entity semantics while enabling higher-order transport and direct field access. Across eleven graphs and multiple node‑classification, link‑prediction, and few‑shot tasks, ICE outperforms all 30 reported supervised and few‑shot comparisons, with core removals and mechanism controls demonstrating the importance of its higher‑order structure and semantic protection.

By Xunkai Li, Xu Wang, Yinlin Zhu, Xiong Yongfu, Yi Liu, Rong-Hua Li, Guoren Wang