arXiv AI By Jingbo Wen, Liang He, Ziqi He

Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation

Read the original on arXiv AI →

arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 28

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.

By Sachin Gopal Wani, Ajay Dholakia, David Ellison
arXiv Machine Learning
22h ago

How Much Can Language Models Gain from Test-Time Computation?

The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.

By Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu
arXiv Computation and Language
Sep 15

Beyond Depth and Width: The Information-Slack Dilemma in Streaming Test-Time Compute

The paper discusses how the same computational task can require different reasoning strategies depending on the order in which evidence arrives, introducing the concept of an "information‑slack dilemma." It argues that early computation may be useful only if its benefits outweigh the costs of later verification, invalidation, and recovery, and proposes a research agenda focused on selective recovery and predictive policies. The authors emphasize evaluating these approaches by separating early‑execution effects, deployment value versus full‑input alternatives, and the added value of predictive policies while considering shared‑resource costs.

By Xiaotian Zhang (Trooly.AI)
arXiv AI
Sep 18

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

When2Think introduces a post‑training framework that dynamically allocates reasoning depth in Large Reasoning Models based on instance difficulty. The method uses Instance‑level Difficulty‑Aware Control (IDAC) to shape rewards with pre‑computed accuracy and token usage statistics, enabling stable, critic‑free optimization without learned reward models. Experiments on mathematical benchmarks show that When2Think improves accuracy‑efficiency trade‑offs, achieving higher Pass@3 scores while reducing token usage compared to baseline models.

By Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak