arXiv Machine Learning

Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport

arXiv:2608. 04234v1 Announce Type: cross Abstract: We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce.

arXiv Computer Vision
Aug 31

Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision

The study investigates whether transferring relational structure from human mental representations to deep neural networks (DNNs) can improve fine‑grained alignment between the two. Using unsupervised Gromov‑Wasserstein optimal transport, the authors show that fine‑tuning pre‑trained DNNs with Relational Knowledge Distillation (RKD) brings the networks close enough to human representations to align at the individual‑object level on a test set of concepts not seen during training. The improvement is driven mainly by a more human‑like global structure of category distances, while local nearest‑neighbor overlap remains largely unchanged.

By Yuria Shimizu, Soh Takahashi, Takato Horii, Masafumi Oizumi
arXiv AI
Aug 20

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

arXiv:2608. 18339v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts during inference.

By Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong
arXiv Machine Learning
Sep 25

GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

GOMA (Graph-Optimized Multimodal Alignment) introduces a dual-embedding approach for multimodal retrieval, separating content embeddings supervised for paired identity from semantic embeddings trained with cross-modal pairs and observed relationships. The method fuses these embeddings, applies semantic agreement to weight graph edges, and uses restart graph propagation to reinforce the initial signal, enabling both single-modality and dual-attribute retrieval. Across six datasets and four tasks, GOMA outperforms 14 external methods on 14 primary metrics, with controlled experiments highlighting the impact of separate supervision, graph regularization, and semantic-guided propagation.

By Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li, Guoren Wang