arXiv AI By Guilin Zhang, Ziqi Tan, Wulan Guo, Kai Zhao, Hongyun Yang, Mei Luo, Qi Ning, Feng Yang

Budget Boundary Effects in Test-Time Mathematical Reasoning

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Sep 4

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

The paper investigates whether large language models’ reasoning traces truly contain early, informative signals or merely reflect budget and difficulty confounds. Using a restart‑controlled truncation probe, the authors compare continuation success rates against from‑scratch restart curves across 178 problem‑model pairs, finding that only one case shows prefix‑limited success and that continuing a model’s own prefix generally outperforms restarting. A difficulty‑controlled test and two generation‑free analyses reveal that early internal signals do not carry outcome information beyond a problem‑difficulty baseline, underscoring the need for proper counterfactual controls.

By Yigit Utku Bulut
arXiv AI
Aug 20

Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation

The paper investigates whether providing candidate solutions during test‑time aggregation improves or harms accuracy compared to a fresh solve that does not use any candidates. Using Qwen3‑4B on AIME‑2025 and HMMT‑2025, the authors find that conditioning on multiple correct candidates boosts accuracy (+0.290), while conditioning on an all‑wrong candidate pool reduces accuracy (−0.123); the effect for a single correct candidate remains unclear. The study also explores structured interventions and placebo controls, but the underlying mechanisms of these effects are not resolved.

By Guiv Farmanfarmaian