Towards Data Science By Abdullahi Dattijo

How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook

Read the original on Towards Data Science →

A practical reproduction of three retrieval baselines, including the crashes, fixes, and score checks that matter for RAG systems. The post How I Reproduced BM25, Dense Retrieval, and SPLADE on a 16GB MacBook appeared first on Towards Data Science .

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma
Towards Data Science
4d ago

I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance.

The article reports on a benchmark that reduced 1,000 Apache Iceberg files to just six, then measured how this consolidation affected query performance across three SQL workloads. It details the methodology and results of the experiment, highlighting changes in execution speed and resource usage. The findings illustrate the trade‑offs between file count and query efficiency in large data systems.

By Thomas Reid