arXiv Computer Vision
3d ago

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.

By Ho-min Park, Byungkon Kang
arXiv Computer Vision
2d ago

Gestalt: Large Multimodal Interplay Model

arXiv:2610.00576v1 Announce Type: new Abstract: In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal...

By Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, Ji-Rong Wen, Yake Wei, Di Hu