arXiv:2606. 28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer.
By Yong Yi Bay, Kathleen A. Yearick
The paper introduces CASE, a dynamic selection combiner that uses a linear gate trained on answer-token hidden states to choose the best candidate answer from a large language model’s samples. It proposes decodability, a leakage‑free metric that predicts when hidden‑state selection will outperform majority voting, achieving a strong correlation (r=0.75) with accuracy gains. CASE improves accuracy by up to 19 points on medium‑difficulty and 16.8 points on hard questions across general and medical LLMs, and its predictive power transfers to unseen scientific domains.
By Zhixiang wang, Ziliang Hong, Ulas Bagci
The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.
By Guiv Farmanfarmaian
arXiv:2606. 12935v1 Announce Type: new Abstract: Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial computational overhead.
By Wenbo Chen, Puheng Li, Mengyang Liu, Weijie Su, Tianpei Xie
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
By Ning Liu
Parallel test-time scaling samples many reasoning traces and majority-votes their answers, improving LLM accuracy but requiring traces to run to completion, incurring substantial computational overhead. We observe that probing partial traces at intermediate checkpoints can extract current answers without disrupting generation, revealing an evolving aggregate vote.