arXiv Computer Vision By Ho-min Park, Byungkon Kang

Mutual Equilibrium: Multimodal Representation Learning through Reciprocal Feedback

Read the original on arXiv Computer Vision →

The paper introduces MEQ, a mutual feedback architecture that iteratively refines two multimodal inputs into coupled embeddings, each embedding incorporating information from the other. By continuously exchanging information between the modalities, the model converges to a fixed point that improves representation quality. Experiments on classification and visual grounding tasks show that MEQ achieves competitive or superior performance compared to concatenation-based baselines, and qualitatively enhances visual grounding when paired with complementary modalities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.