arXiv AI By Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

Read the original on arXiv AI →

MoEMB introduces a mixture‑of‑experts (MoE) approach to scale universal multimodal embeddings (UME) without increasing the size of the output vector or relying on autoregressive decoding. By expanding encoder capacity along the expert axis, MoEMB achieves state‑of‑the‑art performance on MMEB‑V2 and MRMR benchmarks with only 3 B active parameters, outperforming TTE‑based methods that use more than four times as many active parameters and require significantly more compute. The paper also presents the first comprehensive study of adaptive computation for MoE‑based embeddings, exploring training‑time and inference‑time strategies to further improve efficiency for large‑scale retrieval and recommendation systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER is a Unified Multimodal Embedding and Ranking framework that combines contrastive embeddings with Pair‑Aware Discriminative Reasoning to improve universal multimodal retrieval. It replaces item‑wise reflection with pair‑wise comparison of query–candidate pairs, enabling explicit identification of matching and discrepancy evidence. A mutual distillation strategy transfers reliable pairwise preferences between the embedding and ranking components, and UMER achieves state‑of‑the‑art performance on the MMEB‑V2 benchmark while supporting budget‑adjustable inference.

By Libiao Chen, Xiyang Liu, Yanheng Wei, Tao Wang, Zhenyu Tang
arXiv AI
5d ago

TTE-Flash: Accelerating Reasoning-based Multimodal Representations via Think-Then-Embed Tokens

arXiv:2605.16638v2 Announce Type: replace Abstract: Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning. In this pa...

By Jianpeng Cheng, Xian Wu, Jiangfan Zhang, Wentao Bao, Chaitanya Ahuja, Shlok Kumar Mishra, Xuanming Cui, Hanchao Yu, Yang Gao, Fan Xia, Haixing Dai, Frankie Yuan, Zihao Wang, Xiaobing Chen, Qi Guo, Shaodan Zhai, Aashu Singh, Xiangjun Fan, Jun Xiao
arXiv Computer Vision
Sep 16

Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility

The paper introduces the Multi-modal Knowledge Preserving Adapter (MKP-Adapter), an adapter-only approach that enables backward compatible training for multi-modal large language models without updating the backbone. It employs a multi-level preservation loss to maintain embedding geometry and a focal re-weighting strategy to focus on difficult samples. Experiments show strong backward compatibility across image, text, visual document, and video retrieval tasks with minimal latency overhead.

By Jaeseok Byun, Gukyeong Kwon, Han-Kai Hsu, Meher Gitika Karumuri, Zhikang Zhang, Hao Yang, Davide Modolo
arXiv Computation and Language
Aug 27

MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control

MMEmb-R1 is a multimodal embedding framework that enhances reasoning by treating it as a latent variable and selecting beneficial reasoning paths through pair-aware selection and counterfactual intervention. It uses reinforcement learning to invoke reasoning only when necessary, reducing unnecessary computation and latency. On the MMEB-V2 benchmark, MMEmb-R1 achieves a state‑of‑the‑art score of 71.2 with just 4 B parameters.

By Yuchi Wang, Dingkang Yang, Haiyang Yu, Weikang Bian, Jiefeng Long, Xiao Liang, Chao Feng, Hongsheng Li