arXiv AI By Leo Luo, Haining Xie, Siqi Shen, Zhipeng Ma, Rui Ling, Hang Xu, Hefeng Jiang, Dingwei Chen, Yang Li, Peng Chen, Jie Jiang

SIRIUS-SQL: Anchoring Multi-Candidate Text-to-SQL in Execution Feedback

Read the original on arXiv AI →

arXiv:2606. 01246v1 Announce Type: new Abstract: Text-to-SQL on complex schemas is unreliable on a single pass, so recent systems generate multiple SQL candidates and let voting filter out errors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 18

ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models

arXiv:2608. 15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Text-to-SQL inference pipelines.

By Xinmei Huang, Jie Song, Peng Li, Fuxin Jiang, Jing Zhang, Tieying Zhang, Jianjun Chen, Chenming Liu, Tao Yang, Maoyin Liu, Wenda Li, Hong Chen, Cuiping Li
arXiv Computation and Language
Sep 25

ModularSQL: A Runtime Guardrail for the Multiplicity Blind Spot in Text-to-SQL

The paper introduces ModularSQL, a lightweight runtime guardrail designed to detect and correct multiplicity errors—such as missing DISTINCT clauses, inflated aggregates, and Cartesian join explosions—in Text-to-SQL systems. It highlights the Multiplicity Blind Spot (MBS), where standard set-based accuracy metrics fail to capture these errors, and proposes Multiset-EX as a more comprehensive evaluation criterion. Experiments on several models show that ModularSQL can improve multiplicity-aware accuracy while adding minimal computational overhead.

By Tianxin Zhou, Ruixi Lin
arXiv AI
Aug 25

Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries

The paper evaluates using large language model (LLM) juries to review code generated from natural language queries, focusing on MySQL text-to-SQL tasks. It benchmarks 15 open models, selects the top six, and constructs unanimous committees of varying sizes to accept a query only when all members agree. The study finds that single-model judges are inconsistent, while small unanimous committees of strong models can reduce false accepts without discarding many correct queries, and that committee composition significantly influences performance.

By Muhammad Aziz Ullah, Abdul Serwadda
arXiv AI
4d ago

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

The paper introduces a controlled evaluation to disentangle answer coverage, repeatable task advantages, and gains from pre‑execution selection in large‑language‑model (LLM) harnesses. On 386 MATH‑500 tasks, eight generated harnesses and a baseline with nine identical copies were compared over three executions each, revealing that identical programs provide a 2.16‑point repeat‑averaged oracle headroom while generated programs show more repeatable score patterns but mainly expose persistent weaknesses. The study concludes that coverage and repeatability alone cannot justify claims of useful specialization and proposes an evaluation standard for harness diversity that requires task advantages to persist across executions and improve on additional fixed‑program executions under matched inference budgets.

By Ziyang Xu, Haitian Zhong, Hao Zhou, Hao Qin, Chenhan Jin, Te Qi, Shengze Xu, Tieyong Zeng
arXiv Computation and Language
6d ago

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

The paper evaluates a production text‑to‑SQL pipeline that uses an LLM as a judge, finding that the deployed gpt‑4o‑mini judge agrees with human annotators only weakly (Cohen’s kappa 0.04 on a disagreement‑enriched set and 0.42 on a random spot‑check). The authors identify a specific failure mode, GRADE‑HALLUCINATION, responsible for most over‑flags, and demonstrate that a self‑hosted Qwen3.6‑27B model achieves substantially higher agreement (kappa 0.72) at a much lower cost. They also show that ensembling judges does not improve performance, and that their audit method flags a significant portion of out‑of‑domain SQLs as potential issues.

By Haowei Liu, Hsin-Tai Wu, Yi Fang