Hugging Face Trending Papers

Dense Expands, Sparse Anchors: Channel-Asymmetric Query Expansion for Hybrid Retrieval

arXiv AI
Aug 19

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.

By Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
arXiv AI
Sep 3

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search describes DocuSearch, an offline multi‑agent system designed for telecom network operations. The system combines semantic vector search, BM25 full‑text search, and knowledge‑graph neighbor expansion, merges the results via Reciprocal Rank Fusion, and reranks with a cross‑encoder before pruning with Maximal Marginal Relevance. A per‑chunk evaluation loop ensures only grounded answers are returned, achieving Precision@10 of 0.69, Recall@10 of 0.79, and an 89.6% grounding rate—improvements of 15, 16, and 18.4 percentage points over a dense‑only baseline.

By Harish Saragadam, Sudhanshu Sharma, Meghana Pujari
Hugging Face Trending Papers
Sep 8

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Q2D-Web is a new large‑scale benchmark for agentic Retrieval‑Augmented Generation (RAG) systems, featuring a 190 million‑document web corpus and 70 k machine‑reformulated search queries in ten languages. It supplies three sets of relevance judgments—agent citations, production rankings, and a combined set enriched with LLM‑based labels—to evaluate first‑stage retrievers. Experiments on 13 retrievers show consistent ranking across judgment sets but significant variation across domains, languages, and query types, and demonstrate that a carefully sampled sub‑corpus can approximate full‑corpus evaluation with minimal loss in Recall@1000.

arXiv Machine Learning
Sep 24

The Recall Ceiling of LLM Recommendation Reranking

The paper examines LLM-based recommendation rerankers that are often evaluated under an oracle protocol, which guarantees the ground-truth item is present in the scored set. Across Amazon datasets, this protocol overestimates realistic NDCG@10 by 92–95% because realistic retrieval only covers 2–19% of relevant items at K=100, creating a recall ceiling that limits any closed-candidate reranker's top‑k NDCG. The authors find that various optimisation strategies—including prompt engineering, model scaling, sequential models, supervised neural rerankers, LoRA fine‑tuning, hybrid retrieval, score‑aware prompting, and LLM+CF fusion—do not significantly improve over a collaborative‑filtering baseline under realistic retrieval, and they propose a Recall‑Aware Evaluation Protocol (RAEP) to better assess rerankers in low‑recall regimes.

By Zhaohui Wang
arXiv AI
Oct 2

What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation

The paper introduces a streaming-recall benchmark that separates retention and selection decisions for persistent agents. It shows that query‑aware selection boosts recall by 15.5 points when access is fixed, while mixed comparisons inflate gains due to changes in history access. The study finds that under bounded retention, failures stem from eviction rather than ranking errors, and that dense retrieval can outperform lexical retrieval on natural text.

By Juli Huang