arXiv AI By Jinhyung Bae

Spread and Scale: What Determines Whether Test-Time Budget Allocation Pays

Read the original on arXiv AI →

The paper investigates when reallocating a fixed test‑time budget toward harder instances improves solution quality for neural combinatorial optimization solvers. Through pre‑registered experiments on three solvers and two hard‑workload constructions for the traveling salesman problem, it finds that the key deciding factor is the variation in instance difficulty within a workload, not the average difficulty. A budget‑aware policy that first spends part of the budget to gauge instance difficulty recovers most of the potential improvement, though not all, when the cost of this information is included.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

How Much Can Language Models Gain from Test-Time Computation?

The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.

By Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu