arXiv AI

Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

arXiv:2608. 07946v1 Announce Type: cross Abstract: Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean.

arXiv Computation and Language
Aug 31

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction

The paper reports that a model can pass fidelity checks—verifying that extracted values match the source—without actually opening a datasheet, due to a hidden constraint that disables tool use. To address this, the authors log every tool call in an agentic benchmark and develop two instruments: a rule‑based failure‑attribution classifier and a silent‑failure detector that flags runs based solely on which tools were invoked. While the detector shows low false positives on clean extractions and recovers all planted faults, its recall against correct tool usage but incorrect answers remains unmeasured, and a partial causal chamber confirms only a subset of claims, highlighting limitations in physical verification.

By Qing Ye, Meng-Hsuan Lin
arXiv AI
Aug 11

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

arXiv:2608. 09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number.

By Morris Lee
arXiv Computation and Language
3d ago

TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA

TRACE is a system designed to bridge the grounding contract gap in LitTraceQA by combining target-aware retrieval, independent typed evidence localization, multimodal table extraction, and schema-driven table construction. It indexes 27,487 papers using multiple representations while preserving question targets, predicts observation units for tables, and assembles rows with evaluator-compatible key normalization. On the official test set, TRACE achieves a 0.760613 overall score, with high paper F1, evidence F1, and multiple-choice accuracy, though table-row and macro cell performance remain lower.

By Sachin Gupta, Divya Godara
arXiv AI
2d ago

BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

BudgetSchemaBench is a diagnostic tool for evaluating how different schema‑context budgets affect text‑to‑SQL systems. It automatically derives relevance labels from gold SQL, tests four budgets across 80 databases, and compares three schema representations while keeping table rankings fixed. The study shows that increasing the budget improves execution accuracy, especially for lexical retrieval, and that dense retrieval already captures most needed tables at low budgets.

By Chen Shen
arXiv AI
Sep 3

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

The paper introduces ClaimReceipt, a specification and verifier that checks whether a claim in an agent evaluation can be recomputed from retained evidence (sufficiency) and whether the evidence covers the entire experiment set (coverage). Using the CR‑2 verifier on 1,392 historical records, the authors demonstrate accurate reproduction of audit verdicts, non‑redundant field groups, and zero false positives on semantic faults. In a prospective CR‑3 run, the system correctly flags missing receipts and preserves coverage when private evidence is withheld, while adding minimal overhead to inference time and transaction size.

By Peiying Zhu, Sidi Chang
arXiv Machine Learning
Sep 14

What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework

The paper introduces LAST-CQ, a five-agent, training‑free, execution‑grounded framework for Text‑to‑Cypher that evaluates which components of an agentic pipeline contribute most to performance. Experiments on 2,471 live‑database queries across six backbones show that removing correction reduces execution‑BLEU by 3.1–12.3%, while substituting schema‑grounded feedback with raw error strings has negligible impact. Parallel sampling degrades quality by 10–11%, whereas failure detection and retry routing recover 91.7% of initially failed queries, highlighting that simple failure handling is more effective than sophisticated feedback or increased sampling.

By Ioannis Prokopiou, Athanasios Aidinis, Panagiotis-Christos Kyrmpatsos, Pantelis Vikatos