arXiv Machine Learning
23h ago

How Much Can Language Models Gain from Test-Time Computation?

The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.

By Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu
arXiv AI
5d ago

Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability

The paper introduces Budget‑Efficient Thinking (BET), a two‑stage framework that treats adaptive reasoning as a computational investment, aligning solve‑or‑fold decisions with expected return rather than perceived difficulty. BET learns three distinct behaviors: concise short solves for easy queries, early abstention (nice fold) when further reasoning is unlikely to pay off, and allocating sufficient compute (hero call) for hard‑but‑solvable questions. Experiments on seven benchmarks with three base models show BET cuts reasoning tokens by 54% while boosting accuracy by up to 3.2%, and it transfers effectively to scientific QA and logical reasoning tasks.

By Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu, Tingzhao Li, Yiqing Hu, Yumeng Zhao
arXiv AI
2d ago

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

The paper introduces Galahad, a memory layer that stores a transformer language model’s key‑value state for blocks of text, allowing subsequent requests to reuse previously computed attention rather than recomputing it. On seven real‑world datasets, 98.7% of prompt tokens were already read, and with Galahad the model could attend to an entire 97,000‑token corpus, achieving 98–100% recall on a 100‑fact test while reducing inference time and energy consumption dramatically. The approach was validated across 30 models and all runtimes, demonstrating that stateful inference can replace stateless serving without loss of accuracy.

By Sietse Schelpe