arXiv AI By Yanchao Li, Wanhao Liu, Jiaqing Xie, Ben Gao, Yanbo Wang, Tianfan Fu, Yuqiang Li

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Read the original on arXiv AI →

arXiv:2608. 03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv Machine Learning
Aug 11

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv AI
Sep 24

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

The paper investigates why some tasks in the Terminal‑Bench/Frontier‑Bench datasets fail for all agents, distinguishing genuine difficulty from artifacts such as missing context, broken solutions, infrastructure failures, or verifier bypasses. Analyzing 125 all‑fail tasks, only 78 are certified as genuinely unsolved after applying a validity screen; the rest are attributable to broken oracles, infrastructure issues, bypassable verifiers, or insufficient evidence. The study concludes that a zero pass rate does not automatically indicate a hard task and recommends that frontier benchmarks provide evidence for all‑fail tasks before claiming capability gaps.

By Edward Lue Chee Lip, Boden Moraski, Tim Knappe, Lang Xiong, Sarvesh Gharat, Antonio Mari, Ivan Bercovich
arXiv AI
2d ago

From Discovery to Decision: Finite-Budget Recoverability in LLM Voting

The paper studies how voting over multiple large language model (LLM) responses can be optimized under a fixed call budget. It introduces a recoverability threshold that quantifies the gap between discovering a correct answer and ensuring it wins the plurality vote, showing that the candidate set can only grow while the set of reachable winners can only shrink. The authors also present a gold‑free locking certificate that identifies the earliest prefix where all remaining continuations produce the same fixed‑budget output, and demonstrate empirical gains in accuracy and call efficiency through input permutation and exact locking.

By Shaoang Li, Jian Li