arXiv Computation and Language By Aaryan Kapoor, Md Abdullah Al Hafiz Khan

ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval

Read the original on arXiv Computation and Language →

ExecRetrieval is a new benchmark for code‑embedding retrieval that contains 939 Python tasks, each with a verified correct implementation and up to four single‑edit buggy distractors generated mechanically. The dataset allows direct testing of a retriever’s ability to functionally discriminate correct code from near‑clone incorrect code, rather than relying on lexical similarity. Experiments on 23 dense embeddings and BM25 show that while the best system can retrieve the correct code within the top 10 results, it often fails to rank the correct implementation first, with rank‑1 errors dominated by paired buggy variants.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 1

SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue Localization

SpIDER is a dense retrieval method that combines LLM reasoning with graph-based exploration of codebases to locate relevant functions, classes, or files for user queries. It introduces a graph-structured benchmark, SpIDER-Bench, covering multiple programming languages and demonstrates significant recall improvements over traditional BM25 and dense approaches. The method’s graph-based candidate expansion provides auditable structural reasons for each retrieved item while keeping the retrieval budget fixed.

By Shravan Chaudhari, Rahul Thomas Jacob, Jiajun Cao, Shihab Rashid, Mononito Goswami, Christian Bock
arXiv AI
1d ago

Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering

The paper presents a controlled comparison of six retrieval-augmented generation (RAG) strategies for scientific question answering on a large arXiv corpus. All pipelines use the same LLM generator and evaluation protocol, differing only in retrieval design—ranging from classic dense retrieval to late‑interaction methods like ColBERTv2. The authors also release a synthetic question dataset and code to enable reproducible, large‑scale evaluation of RAG trade‑offs.

By Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos