arXiv AI By Wenxiao Fan, Jingling Fu, Luohang Liu, Xinyuan Shan, Lichen Ma, Yu He, Junshi Huang, Yan Li, Kan Li

Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings

Read the original on arXiv AI →

The paper investigates whether reasoning always benefits universal multimodal embeddings (UMEs). By comparing the discriminative and reasoning-driven branches of UME-R1, the authors find that while reasoning improves positive similarity in 56.6% of cases, it also creates 15.7% false-helpful instances where hard negatives are drawn closer. Diagnostic analyses reveal that reasoning often de‑condenses retrieved neighborhoods and that chain‑of‑thought tokens encode evidence common to both positives and hard negatives. Based on these insights, the authors introduce SURE, a utility router that boosts UME-R1‑7B by 1.5 points and consistently improves other embedding models on MMEB‑V2 without retraining or extra VLM passes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER is a Unified Multimodal Embedding and Ranking framework that combines contrastive embeddings with Pair‑Aware Discriminative Reasoning to improve universal multimodal retrieval. It replaces item‑wise reflection with pair‑wise comparison of query–candidate pairs, enabling explicit identification of matching and discrepancy evidence. A mutual distillation strategy transfers reliable pairwise preferences between the embedding and ranking components, and UMER achieves state‑of‑the‑art performance on the MMEB‑V2 benchmark while supporting budget‑adjustable inference.

By Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang
arXiv AI
5d ago

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

arXiv:2605.16638v2 Announce Type: replace Abstract: Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this pa...

By Jianpeng Cheng, Xian Wu, Jiangfan Zhang, Wentao Bao, Chaitanya Ahuja, Shlok Kumar Mishra, Xuanming Cui, Hanchao Yu, Yang Gao, Fan Xia, Haixing Dai, Frankie Yuan, Zihao Wang, Xiaobing Chen, Qi Guo, Shaodan Zhai, Aashu Singh, Xiangjun Fan, Jun Xiao
arXiv Computation and Language
Aug 27

MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control

MMEmb-R1 is a multimodal embedding framework that enhances reasoning by treating it as a latent variable and selecting beneficial reasoning paths through pair-aware selection and counterfactual intervention. It uses reinforcement learning to invoke reasoning only when necessary, reducing unnecessary computation and latency. On the MMEB-V2 benchmark, MMEmb-R1 achieves a state‑of‑the‑art score of 71.2 with just 4 B parameters.

By Yuchi Wang, Dingkang Yang, Haiyang Yu, Weikang Bian, Jiefeng Long, Xiao Liang, Chao Feng, Hongsheng Li
arXiv Computer Vision
Sep 1

Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings

arXiv:2604.22280v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that...

By Peixi Wu, Ke Mei, Feipeng Ma, Bosong Chai, Zhibin Lan, Chenxi Zhao, Shannan Yan, Jie Chen, Zhangchi Hu, Yansong Peng, Bo Lin, Junjie Zhou, Dacheng Yin, Tianyi Wang, Fengyun Rao, Jing Lyu, Hebei Li, Xiaoyan Sun