arXiv AI

Access Paths for Efficient Ordering with Large Language Models

arXiv AI
Jun 30

SemJoin: Semantic Join Optimization

arXiv:2606. 29532v1 Announce Type: cross Abstract: Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and analysis.

By Christopher Gou, Aditya Banerjee, Jiaxuan Wang, Chunwei Liu
arXiv AI
Jul 28

Kalypso: Relational LLM Serving

arXiv:2607. 23815v1 Announce Type: cross Abstract: Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data.

By Hojae Son, Md Ashraful Islam, Huy Gia Cao, Hui Guan, Marco Serafini
Towards Data Science
Sep 28

Guided Merge Sort : An Optimized Sorting that Picks the Best from Ordinary and Multi-Way Merge Sort Algorithms

The article introduces Guided Merge Sort, an optimized sorting technique that combines elements of ordinary merge sort and multi‑way merge sort. It highlights how the use of the "goto" operator becomes essential in this approach. The post explains the algorithm’s design and its potential advantages over traditional methods.

By Tigran Hayrapetyan
arXiv Machine Learning
Sep 1

QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph Generation

QueryGraph is a system that transforms natural language queries into structured graphs for reliable multi-tool execution. It employs a deterministic planner that uses depth-first search to resolve tool dependencies and combine results, improving reliability over traditional keyword searches. The approach works well even with smaller or locally hosted large language models, achieving high accuracy in multi-step, cross-tool queries.

By Aishwarya Chakravarthy, Vidhi Kulkarni, Duen Horng Chau
arXiv AI
Sep 2

Replacing Training with Memory: Listwise Selection for Text-to-SQL

The paper proposes a fine‑tuning‑free listwise selector for Text‑to‑SQL systems that replaces traditional learning objectives with inference‑time strategies. It introduces reusable structured memories (MaP‑SQL) that encode mappings from natural language to schema elements, SQL operations, and expected outputs, and uses these memories to evaluate candidate queries. To reduce positional bias, the method aggregates rankings across multiple input permutations, optimizing inference cost through execution results and pointwise scoring. The approach achieves higher selection accuracy, fewer unnecessary comparisons, and outperforms the prior state‑of‑the‑art R^3‑SQL on the BIRD‑dev benchmark while using fewer tokens.

By Yeonseok Jeong, Soyoung Yoon, Seongjun Lee, Seung-won Hwang