arXiv AI By Yong Yi Bay, Kathleen A. Yearick

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

Read the original on arXiv AI →

arXiv:2606. 28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 3

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

arXiv:2606. 02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assumption that longer reasoning is consistently beneficial remains under-examined.

By Simone Caldarella, Davide Talon, Rahaf Aljundi, Elisa Ricci, Massimiliano Mancini
arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv AI
2d ago

From Discovery to Decision: Finite-Budget Recoverability in LLM Voting

The paper studies how voting over multiple large language model (LLM) responses can be optimized under a fixed call budget. It introduces a recoverability threshold that quantifies the gap between discovering a correct answer and ensuring it wins the plurality vote, showing that the candidate set can only grow while the set of reachable winners can only shrink. The authors also present a gold‑free locking certificate that identifies the earliest prefix where all remaining continuations produce the same fixed‑budget output, and demonstrate empirical gains in accuracy and call efficiency through input permutation and exact locking.

By Shaoang Li, Jian Li
arXiv AI
Jul 17

Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

arXiv:2607. 14552v1 Announce Type: cross Abstract: A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep those that reach the correct final answer, and fine-tune on the survivors.

By Jungseob Lee, Seungyoon Lee, Suhyune Son, Dongyub Jude Lee, Sungbin Han, Sugyeong Eo, Heuiseok Lim