arXiv AI

Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.

arXiv AI
Aug 7

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

arXiv:2608. 05643v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity.

By Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Lena Trigg, Ali Subhan, Muhammad Ali, Dean F. Hougen
arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv AI
1d ago

Budget Boundary Effects in Test-Time Mathematical Reasoning

arXiv:2609.38699v1 Announce Type: new Abstract: A cumulative token cap can fall inside a mathematical derivation, forcing a test-time controller to choose between stopping at the cap (strict) and all...

By Guilin Zhang, Ziqi Tan, Wulan Guo, Kai Zhao, Hongyun Yang, Mei Luo, Qi Ning, Feng Yang
arXiv Machine Learning
Sep 25

Return or Revise? Learning When Revision Helps Retrieval-Augmented QA

The paper investigates when it is better to return an existing draft answer or revise it using retrieved evidence in retrieval‑augmented QA systems. By grading both the draft and its candidate revision with the same correctness judge, the authors define a paired effect called recoverability and train policies to predict it before revision. Experiments on 25,870 open‑domain questions show that a recoverability‑based scorer outperforms a draft‑correctness scorer across multiple Llama setups, improving accuracy–revision trade‑offs and closing a significant portion of the oracle gap, though it still applies harmful revisions in a substantial fraction of cases.

By Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup, Grant Erdmann
arXiv AI
Sep 18

LLM-as-an-Improver: Turning Verification into Better Candidates

The paper introduces LLM-as-an-Improver, a method that uses verification feedback to enhance the candidate set in verifier-based selection. It proposes Verify–Repair–Reselect (VRR), which keeps the initial winner, generates three complementary alternatives (repaired versions of the winner and runner‑up, and a new approach), filters invalid or duplicate candidates, and then reselects the final answer. Experiments on code‑generation and reasoning benchmarks show that VRR outperforms fixed‑pool selection and can recover correct solutions even when the initial pool is entirely wrong.

By Akiyoshi Tomihari, Yuma Ichikawa
arXiv AI
Sep 1

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.

By Qiancheng Zhou, Ruizhe Li
arXiv AI
Aug 3

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.

By Penglin Zhu, Jungang Xu