arXiv AI

Semantic Early-Stopping for Iterative LLM Agent Loops

arXiv:2606. 27009v1 Announce Type: new Abstract: Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a fixed iteration cap (max_iterations).

arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv Machine Learning
1d ago

How Much Can Language Models Gain from Test-Time Computation?

The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.

By Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, Qinglong Zheng, Zhanghan Ni, Ge Liu
arXiv AI
Sep 4

</think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination

The paper investigates a training‑free early‑exit technique that inserts an end‑of‑think (EoT) token to terminate chain‑of‑thought (CoT) reasoning in large reasoning models. It finds that the injected EoT often fails to cleanly switch the model from reasoning to answering, leading to continued reasoning‑like generation—termed spurious CoT termination—whose length scales with the amount of reasoning saved. By increasing attention to the EoT token through Exit‑token Attention Biasing (EAB), the authors reduce spurious termination and shorten the answering phase across multiple models and benchmarks.

By Seunghee Koh, Sungjae Choi, Minchan Kwon, Sunghyun Baek, Junmo Kim
arXiv Computation and Language
Sep 15

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...

By Kaiyuan Liu, Qiuyang Mang, Bo Peng, Wenhao Chai, Hanchen Li, Shreyas Pimpalgaonkar, Luke Zettlemoyer, Alex Dimakis, Alvin Cheung
arXiv AI
Aug 17

The Metacognitive Bottleneck: Japanese Riddles Reveal Fundamental Limits of Machine Insight and Self-Evaluation in Reasoning AI

arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.

By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen
arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra
Hugging Face Trending Papers
Jun 10

Teaching Diffusion to Speculate Left-to-Right

Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.