arXiv Computation and Language By Ty Chermsirivatana, John MacCormick

Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

Read the original on arXiv Computation and Language →

The paper investigates how increasing inference-time computation—via wider beam search or sample‑plus‑vote—affects performance on grammar‑constrained text‑to‑SQL tasks for small language models. Using the Qwen2.5‑Instruct family (0.5B–7B parameters) on the Spider benchmark, the authors find that larger models consistently outperform higher inference compute on the same model size, and that beam search yields better accuracy than sample‑plus‑vote under matched budgets. These results suggest that, unlike unconstrained settings, scaling inference compute does not compensate for smaller model size when strict grammar constraints are applied.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 17

Register Bias in Complexity-Based Large Language Model Routing

The paper examines how large language model (LLM) services route queries to models of varying size based on a cheap complexity estimate. It finds that this routing is not register neutral: queries written in non‑standard English registers (e.g., African American English or second‑language English) are systematically assigned to lower‑capacity models because they appear shorter due to omitted function words. Experiments on 37,704 learner sentence pairs and a controlled corpus show that this bias leads to significantly lower accuracy across all model tiers, including the highest‑capacity cloud models, while the routing decision itself adds little marginal cost.

By Simran Koul