Hugging Face Trending Papers

SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval

Read the original on Hugging Face Trending Papers →

With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant documents from multilingual text collections using a single-language query.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 28

MIMO: Multilingual Information Retrieval via Monolingual Objectives

MIMO: Multilingual Information Retrieval via Monolingual Objectives proposes a two‑stage framework that first aligns a student model to a stable English semantic space using knowledge distillation, then jointly optimizes distillation and cross‑lingual contrastive learning to improve retrieval discrimination while preserving alignment. The approach addresses language clustering and the trade‑off between cross‑lingual alignment and embedding uniformity, outperforming existing cross‑lingual training baselines on both multilingual and multi‑monolingual benchmarks. MIMO also remains competitive with larger off‑the‑shelf models and its alignment‑uniformity analysis clarifies the distinct roles of the two loss components. whyItMatters":"The study demonstrates a practical method to enhance multilingual information retrieval performance by balancing alignment and uniformity, which is crucial for real‑world search environments where queries and documents span multiple languages."

By Youngjoon Jang, Seongtae Hong, Heuiseok Lim
arXiv Computation and Language
Aug 27

Align Then Adapt: Label-Efficient Adapter Learning for Asymmetric Dense Retrieval

The paper introduces Efficient Retrieval Adapter (ERA), a query‑side adapter framework that enables dense retrieval systems to adapt to asymmetric query–document scenarios without re‑indexing. ERA first aligns the embedding spaces of a powerful query embedder and a lightweight document embedder using unlabeled documents, then fine‑tunes the aligned query representation with a small set of labeled query‑document pairs. In experiments on 126 MAIR retrieval tasks across six domains, ERA boosts average nDCG@10 by up to 8.2 points in symmetric settings and over 12 points in asymmetric settings while requiring far fewer labels than fully supervised adapter training.

By Seiji Maekawa, Moin Aminnaseri, Pouya Pezeshkpour, Estevam Hruschka
arXiv AI
Aug 19

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.

By Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
arXiv AI
Jun 15

Succeeding at Scale: Enterprise Retrieval Benchmark Construction and Index-Preserving Query Adaptation for Multi-Tenant Search

arXiv:2601. 04646v4 Announce Type: replace-cross Abstract: Large-scale multi-tenant retrieval systems generate extensive query logs but lack curated relevance labels for effective domain adaptation, resulting in substantial underutilized "dark data.

By Prateek Jain, Shabari S Nair, Ritesh Goru, Prakhar Agarwal, Ajay Yadav, Yoga Sri Varshan Varadharajan, Constantine Caramanis