The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.
By Yigit Utku Bulut
arXiv:2608. 04160v1 Announce Type: cross Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable.
By Ankit Goyal, Jaideep Ray
arXiv:2608. 03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency.
By Francesca Carlon, Vincent Ginis, Andres Algaba
The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.
By Guiv Farmanfarmaian
arXiv:2608. 07303v1 Announce Type: new Abstract: Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong.
By Guilin Zhang, Kai Zhao
arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.
By Sajib Acharjee Dip, Dawei Zhou, Liqing Zhang
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
By Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin
arXiv:2608. 12150v1 Announce Type: new Abstract: Standard evaluation of large language models assumes stable model rankings across inference conditions.
By Rodrigo Guedes de Souza, Alison R. Panisson
arXiv:2608. 07341v1 Announce Type: cross Abstract: Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized.
By Ruijie Hou, Yueyang Jiao, Zhao Wang, Yingming Li
arXiv:2608. 06362v1 Announce Type: cross Abstract: Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time.
By Boning Li, Yu Chen, Longbo Huang
arXiv:2609.36721v1 Announce Type: new
Abstract: Fixed-budget adaptation from heterogeneous data sources requires deciding not only how much data to use, but how much exposure each source receives. Si...
By Yupeng Chang, Yuan Wu
The paper introduces a calibrated instrument for rigorously measuring how inference optimizations—such as quantization, early‑exit, and speculative decoding—affect the output quality of large language models. It uses a formally calibrated LLM judge that verifies no systematic bias between statistically equivalent outputs and includes a null condition to ensure measured differences are zero. Applying this method, the authors find that a 4‑bit model is indistinguishable from its 16‑bit counterpart, while 3‑bit quantization and early‑exit techniques incur measurable quality losses that vary by language and task, and that token‑certainty‑based acceptance rules cannot reliably identify impactful errors.
By Jerry Kaplan