When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
arXiv:2606. 28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer.
The paper studies how voting over multiple large language model (LLM) responses can be optimized under a fixed call budget. It introduces a recoverability threshold that quantifies the gap between discovering a correct answer and ensuring it wins the plurality vote, showing that the candidate set can only grow while the set of reachable winners can only shrink. The authors also present a gold‑free locking certificate that identifies the earliest prefix where all remaining continuations produce the same fixed‑budget output, and demonstrate empirical gains in accuracy and call efficiency through input permutation and exact locking.
arXiv:2606. 28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer.
The paper introduces CASE, a dynamic selection combiner that uses a linear gate trained on answer-token hidden states to choose the best candidate answer from a large language model’s samples. It proposes decodability, a leakage‑free metric that predicts when hidden‑state selection will outperform majority voting, achieving a strong correlation (r=0.75) with accuracy gains. CASE improves accuracy by up to 19 points on medium‑difficulty and 16.8 points on hard questions across general and medical LLMs, and its predictive power transfers to unseen scientific domains.
The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.
arXiv:2606. 12935v1 Announce Type: new Abstract: Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial computational overhead.
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial computational overhead. We observe that probing partial traces at intermediate checkpoints can extract current answers without disrupting generation, revealing an evolving aggregate vote.
ReSolve is a training‑free inference method that reuses candidate reasoning by selectively moderating generative outputs. It examines existing derivations when candidates disagree or lack a parseable answer, then incorporates new solutions into a bounded loop. On 130 competition‑mathematics problems, ReSolve achieves 100 and 99 correct answers with significantly fewer tokens than eight‑sample self‑consistency, while a controlled ablation shows that visible derivations improve accuracy.
arXiv:2608. 15445v1 Announce Type: new Abstract: When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization.
The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.
arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
arXiv:2608. 11403v1 Announce Type: new Abstract: Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plurality answer.