Query Circuits: Explaining How Language Models Answer User Prompts
arXiv:2509. 24808v2 Announce Type: replace Abstract: Explaining why a language model produces a particular output requires local, input-level explanations.
The paper investigates how increasing inference-time computation—via wider beam search or sample‑plus‑vote—affects performance on grammar‑constrained text‑to‑SQL tasks for small language models. Using the Qwen2.5‑Instruct family (0.5B–7B parameters) on the Spider benchmark, the authors find that larger models consistently outperform higher inference compute on the same model size, and that beam search yields better accuracy than sample‑plus‑vote under matched budgets. These results suggest that, unlike unconstrained settings, scaling inference compute does not compensate for smaller model size when strict grammar constraints are applied.
arXiv:2509. 24808v2 Announce Type: replace Abstract: Explaining why a language model produces a particular output requires local, input-level explanations.
arXiv:2608. 10137v1 Announce Type: cross Abstract: Grammar Constrained Decoding (GCD) forces Language Models (LMs) to produce syntactically valid outputs by masking out non-conforming tokens at each step.
The paper examines how large language model (LLM) services route queries to models of varying size based on a cheap complexity estimate. It finds that this routing is not register neutral: queries written in non‑standard English registers (e.g., African American English or second‑language English) are systematically assigned to lower‑capacity models because they appear shorter due to omitted function words. Experiments on 37,704 learner sentence pairs and a controlled corpus show that this bias leads to significantly lower accuracy across all model tiers, including the highest‑capacity cloud models, while the routing decision itself adds little marginal cost.
arXiv:2606. 30473v1 Announce Type: cross Abstract: We study retrieval over catalogs of structured metadata, where each record is a small schema whose fields answer different kinds of query.
arXiv:2606. 26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run.
arXiv:2608. 12426v1 Announce Type: new Abstract: Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasoning structure, safety boundaries, output schemas.
arXiv:2609.13486v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning decoder-only large language models (LLMs) for retrieval yields strong first-stage retrievers, with effectivene...
arXiv:2609.39368v1 Announce Type: new Abstract: A common formalism for constraining the output of autoregressive text generation models involves lexical constraints, words or phrases which are requir...
arXiv:2606. 24083v1 Announce Type: cross Abstract: "Talk short.
arXiv:2607. 09438v1 Announce Type: cross Abstract: Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear.
arXiv:2511. 12309v2 Announce Type: replace-cross Abstract: Self-consistency (SC) is a widely used test-time inference technique for improving performance in chain-of-thought reasoning.
arXiv:2609.38006v1 Announce Type: new Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable tha...