Hugging Face Trending Papers

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval.

arXiv AI
Sep 4

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

CORE improves compositional reasoning in multimodal language models by distilling a cross‑attentive reranker’s fine‑grained judgments into the embedding model. It generates candidate lists across five compositional matching levels and trains with a Rank‑KL objective to replicate the reranker’s ranking. Experiments on COLA, SUGARCREPE++, and NEGBENCH show CORE‑RERANKER‑8B outperforms Jina‑Reranker by 10.7 points, while CORE‑EMBED‑8B achieves the best overall average among evaluated embeddings, with gains also transferring to the MCMR benchmark without harming COCO or Flickr30K retrieval.

By Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu
arXiv AI
Sep 25

Is Reasoning Always Useful? Rethinking Reasoning Utility in Universal Multimodal Embeddings

The paper investigates whether reasoning always benefits universal multimodal embeddings (UMEs). By comparing the discriminative and reasoning-driven branches of UME-R1, the authors find that while reasoning improves positive similarity in 56.6% of cases, it also creates 15.7% false-helpful instances where hard negatives are drawn closer. Diagnostic analyses reveal that reasoning often de‑condenses retrieved neighborhoods and that chain‑of‑thought tokens encode evidence common to both positives and hard negatives. Based on these insights, the authors introduce SURE, a utility router that boosts UME-R1‑7B by 1.5 points and consistently improves other embedding models on MMEB‑V2 without retraining or extra VLM passes.

By Wenxiao Fan, Jingling Fu, Luohang Liu, Xinyuan Shan, Lichen Ma, Yu He, Junshi Huang, Yan Li, Kan Li
arXiv Machine Learning
Sep 2

Can LLMs Use Relational Transformer Embeddings?

The paper investigates whether large language models (LLMs) can leverage frozen relational‑transformer embeddings by injecting them as soft tokens. Using a learned MLP projection and LoRA adaptation, the authors fine‑tune Qwen3.5‑4B on chain‑of‑thought reasoning traces and group‑based reinforcement learning, then evaluate on ten binary classification tasks across six RelBench databases. The hybrid approach consistently underperforms the standalone relational transformer, showing sensitivity to serialization format, token budget, and RL stability, leading the authors to conclude that stronger alignment objectives and schema‑aware design are needed for reliable relational prediction.

By Francisco Galuppo Azevedo, Clarissa Lima Loures
arXiv AI
Sep 10

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

MoEMB introduces a mixture‑of‑experts (MoE) approach to scale universal multimodal embeddings (UME) without increasing the size of the output vector or relying on autoregressive decoding. By expanding encoder capacity along the expert axis, MoEMB achieves state‑of‑the‑art performance on MMEB‑V2 and MRMR benchmarks with only 3 B active parameters, outperforming TTE‑based methods that use more than four times as many active parameters and require significantly more compute. The paper also presents the first comprehensive study of adaptive computation for MoE‑based embeddings, exploring training‑time and inference‑time strategies to further improve efficiency for large‑scale retrieval and recommendation systems.

By Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng
arXiv Machine Learning
Sep 7

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation

The paper presents a two-level framework for scalable trade‑up recommendation. Level 1 distills large‑language‑model reasoning into a compact, non‑generative student that classifies product pairs using only precomputed embeddings, achieving high AUC on a benchmark. Level 2 applies product‑type test‑time training to fine‑tune lightweight adapters, further improving performance while keeping inference fast and inexpensive.

By Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi
arXiv AI
Aug 19

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.

By Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
arXiv AI
Jun 30

SemJoin: Semantic Join Optimization

arXiv:2606. 29532v1 Announce Type: cross Abstract: Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis.

By Christopher Gou, Aditya Banerjee, Jiaxuan Wang, Chunwei Liu