arXiv AI

TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data

arXiv:2607. 03926v1 Announce Type: cross Abstract: Synthetic tabular data support use cases like data sharing, model development under access restrictions, and rapid prototyping of analytical workflows.

arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv AI
2d ago

BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

BudgetSchemaBench is a diagnostic tool for evaluating how different schema‑context budgets affect text‑to‑SQL systems. It automatically derives relevance labels from gold SQL, tests four budgets across 80 databases, and compares three schema representations while keeping table rankings fixed. The study shows that increasing the budget improves execution accuracy, especially for lexical retrieval, and that dense retrieval already captures most needed tables at low budgets.

By Chen Shen
arXiv Machine Learning
Jul 7

Identifiability of Relational Queries in Multi-View Pretraining

arXiv:2607. 04735v1 Announce Type: cross Abstract: When data sources are integrated through a shared interface, a downstream query may or may not be determined by what the interface exposes: two globally consistent worlds can agree on every shared attribute yet disagree on the query answer.

By Ratan Bahadur Thapa, Daniel Hern\'andez
Hugging Face Trending Papers
Jul 14

Finding the Right Tables and Columns: A Benchmark and Corpus-Adaptive Embeddings for SQL Schema Retrieval

Retrieval in the SQL setting has largely been studied as the task of finding, within a large collection of SQL statements, the statement that answers a natural-language question. At scale, however, a more fundamental retrieval problem precedes generation: schema retrieval, identifying the tables and columns a question requires in a database that may contain thousands of them, far more than fit in a model's context.