Towards Data Science

Guided Merge Sort : An Optimized Sorting that Picks the Best from Ordinary and Multi-Way Merge Sort Algorithms

The article introduces Guided Merge Sort, an optimized sorting technique that combines elements of ordinary merge sort and multi‑way merge sort. It highlights how the use of the "goto" operator becomes essential in this approach. The post explains the algorithm’s design and its potential advantages over traditional methods.

arXiv AI
Sep 15

Access Paths for Efficient Ordering with Large Language Models

arXiv:2509.00303v4 Announce Type: replace-cross Abstract: In this work, we present the \texttt{LLM ORDER BY} semantic operator as a logical abstraction and conduct a systematic study of its physical...

By Fuheng Zhao, Jiayue Chen, Yiming Pan, Tahseen Rabbani, Sohaib, Divyakant Agrawal, Amr El Abbadi, Paritosh Aggarwal, Anupam Datta, Dimitris Tsirogiannis
Towards Data Science
Jul 6

Stop Ranking Agent Configs by Average Score

Best-worst comparisons, MaxDiff-style judging, and Plackett-Luce utility scores give agent teams a cleaner way to decide which configs to ship, prune, and route toward next. The post Stop Ranking Agent Configs by Average Score appeared first on Towards Data Science .

By Doster Esh
Towards Data Science
Jul 22

Loop Engineering for RAG Generation: Iterate top-k One at a Time

Enterprise Document Intelligence [Vol. 1 #8bis] - Two regimes for sending retrieved candidates to the generation brick, the sufficiency signal that picks between them, and the per-question type dispatch that makes it cheap The post Loop Engineering for RAG Generation: Iterate top-k One at a Time appeared first on Towards Data Science .

By Kezhan Shi
Towards Data Science
Aug 25

Recursive CTEs: SQL’s Hidden Graph Traversal Engine

The article titled "Recursive CTEs: SQL’s Hidden Graph Traversal Engine" offers a practical guide for working with hierarchies in SQL. It explains how to navigate these structures, find routes, detect cycles, and calculate degrees of separation using recursive common table expressions. The guide is aimed at readers looking to leverage SQL’s built‑in graph traversal capabilities.

By Thomas Reid
Towards Data Science
Aug 31

Why RAG Complexity Should Be Earned

The article outlines a framework for constructing Retrieval-Augmented Generation (RAG) pipelines that progressively add complexity as needed to address observed failure modes. It begins with basic lexical and hybrid search techniques, then incorporates reranking and agentic information‑seeking strategies to improve performance. The approach emphasizes that more sophisticated components should only be introduced when simpler methods prove insufficient.

By Tahreem Rasul
Towards Data Science
Aug 26

How Does a RAG Reranker Really Work?

The article "How Does a RAG Reranker Really Work?" explores the inner workings of Retrieval-Augmented Generation (RAG) rerankers, focusing on how data scientists explain the model’s operations behind the scenes. It discusses the impact of these insights on architecture decisions within enterprise document intelligence, specifically in the context of Enterprise Document Intelligence Vol.1 #2D. The piece highlights the importance of transparent model explanations for effective enterprise RAG implementation.

By Kezhan Shi
arXiv Machine Learning
Sep 18

A Table-Free Index for Tapered Memoization Grids: Compact Out-of-Core Evaluation of Functions of Sorted Arguments

The paper presents a table‑free index for tapered memoization grids, enabling compact out‑of‑core evaluation of functions that depend on sorted arguments. By showing that the grid’s key set corresponds to multiset combinations, the authors derive a closed‑form O(d) ranking and unranking scheme that removes the need for large preprocessing tables and allows order‑free parallel construction. The resulting values‑only flat array uses significantly less memory than hash‑map memoization, offers faster query times once cache limits are exceeded, and remains operable with memory‑mapped storage beyond RAM.

By Tamal Maharaj