Multimodal representation learning seeks shared representations for cross-modal retrieval and knowledge transfer. Hub-based binding reduces pairwise supervision costs, but separate hub connections can...
arXiv:2506. 08774v2 Announce Type: replace-cross Abstract: Different machine learning models can represent the same underlying concept in different ways.
By Fan Xu, Luis A. Leiva
arXiv:2606. 12362v1 Announce Type: cross Abstract: We study multimodal learning under missing modalities, with particular motivation from bioscience applications in which heterogeneous modalities are often only partially available when decisions need to be made.
By Hui Wang, Tianyu Ren, Joseph Butler, Christopher Baker, Karen Rafferty, Simon McDade
arXiv:2607. 16789v1 Announce Type: new Abstract: Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities.
By Sana Tonekaboni, Viktoria Schuster, Caroline Uhler
Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, prior approaches often overlook cases where images appear visually alike yet differ in attributes, potentially undermining both multimodal feature fusion and similarity modeling.
arXiv:2602. 17395v2 Announce Type: replace-cross Abstract: Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known classes.
By Lorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. Bagdanov
arXiv:2505. 19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities.
By Sanghyuk Chun, Olga Russakovsky
GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.
By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang
arXiv:2607. 15687v1 Announce Type: new Abstract: Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs.
By Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang
ICE (Interaction-aware Clifford Encoder) is a multimodal graph foundation model that uses a node-indexed Clifford latent field to encode topology, text, and images into explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods, preserving entity semantics while enabling higher-order transport and direct field access. Across eleven graphs and multiple node‑classification, link‑prediction, and few‑shot tasks, ICE outperforms all 30 reported supervised and few‑shot comparisons, with core removals and mechanism controls demonstrating the importance of its higher‑order structure and semantic protection.
By Xunkai Li, Xu Wang, Yinlin Zhu, Xiong Yongfu, Yi Liu, Rong-Hua Li, Guoren Wang
arXiv:2606. 10504v1 Announce Type: new Abstract: Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e.
By Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang
arXiv:2607. 08839v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference.
By Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji