arXiv AI By Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz, Amine Mhedhbi

The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

Read the original on arXiv AI →

The paper examines the role of schema linking in Text-to-SQL systems and finds that recent large language models can effectively use relevant schema elements even when many irrelevant ones are present. Consequently, the authors eliminate schema linking when the entire schema fits within the model’s context window, instead employing augmentation, selection, and correction techniques to enhance accuracy. Their approach achieves first place on the BIRD benchmark with a 71.83% accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB

The paper introduces FlockMTL, an extension for database management systems that deeply integrates large language models and retrieval‑augmented generation into DuckDB. It provides model‑driven scalar and aggregate functions, cost‑based optimizations like batching and caching, and new SQL DDL abstractions (PROMPT and MODEL) to treat LLMs as first‑class schema objects. These features aim to simplify the development of knowledge‑intensive analytical applications by reducing the effort required to orchestrate heterogeneous data systems and manage LLM context.

By Anas Dorbani, Sunny Yasser, Jimmy Lin, Amine Mhedhbi
Hugging Face Trending Papers
Jul 14

Finding the Right Tables and Columns: A Benchmark and Corpus-Adaptive Embeddings for SQL Schema Retrieval

Retrieval in the SQL setting has largely been studied as the task of finding, within a large collection of SQL statements, the statement that answers a natural-language question. At scale, however, a more fundamental retrieval problem precedes generation: schema retrieval, identifying the tables and columns a question requires in a database that may contain thousands of them, far more than fit in a model's context.