Hugging Face Trending Papers

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Read the original on Hugging Face Trending Papers →

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv Machine Learning
Aug 11

Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing

arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv AI
Aug 11

Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline

arXiv:2608. 09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number.

By Morris Lee
arXiv Computation and Language
Aug 25

Most of the LLM routing gap is task type

The paper investigates why large‑language‑model (LLM) routers—systems that select the best model for each query—often fail to outperform a single best model. By evaluating 14 models on 294 questions across seven task types and three languages, the authors find that a simple static mapping of task type to model improves 21 of the 29 questions that routing could potentially solve, and that learned routers do not significantly exceed this performance. The study highlights that most routing gains stem from task‑type specialization rather than complex learned decision rules.

By Janghoon Lee