Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.
arXiv:2608. 03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability.
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.
The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.
arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.
The paper investigates why some tasks in the Terminal‑Bench/Frontier‑Bench datasets fail for all agents, distinguishing genuine difficulty from artifacts such as missing context, broken solutions, infrastructure failures, or verifier bypasses. Analyzing 125 all‑fail tasks, only 78 are certified as genuinely unsolved after applying a validity screen; the rest are attributable to broken oracles, infrastructure issues, bypassable verifiers, or insufficient evidence. The study concludes that a zero pass rate does not automatically indicate a hard task and recommends that frontier benchmarks provide evidence for all‑fail tasks before claiming capability gaps.
arXiv:2608.23023v2 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall...
The paper studies how voting over multiple large language model (LLM) responses can be optimized under a fixed call budget. It introduces a recoverability threshold that quantifies the gap between discovering a correct answer and ensuring it wins the plurality vote, showing that the candidate set can only grow while the set of reachable winners can only shrink. The authors also present a gold‑free locking certificate that identifies the earliest prefix where all remaining continuations produce the same fixed‑budget output, and demonstrate empirical gains in accuracy and call efficiency through input permutation and exact locking.
arXiv:2609.05637v2 Announce Type: replace Abstract: A popular way to improve Retrieval-Augmented Generation (RAG) is to rewrite the user's question into several variants and search with all of them....
The paper investigates why large‑language‑model (LLM) routers—systems that select the best model for each query—often fail to outperform a single best model. By evaluating 14 models on 294 questions across seven task types and three languages, the authors find that a simple static mapping of task type to model improves 21 of the 29 questions that routing could potentially solve, and that learned routers do not significantly exceed this performance. The study highlights that most routing gains stem from task‑type specialization rather than complex learned decision rules.
arXiv:2607. 23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which shows correctness.
arXiv:2607. 17240v1 Announce Type: new Abstract: When does a committed intermediate stage in an LLM reasoning pipeline earn its cost?
arXiv:2608. 09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number.
arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.