arXiv Computation and Language

Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

arXiv AI
1d ago

Learning to Ask: Information Acquisition for SLM-LLM Collaboration, under a budget

The paper proposes a new framework for collaboration between a small language model (SLM) and a large language model (LLM) that treats the interaction as an information acquisition problem under an API budget constraint. Instead of delegating reasoning tasks, the SLM remains the primary reasoner and selectively queries the LLM advisor with targeted questions, using a three-stage RLVR approach to decide when to call the advisor, how to phrase queries, and how to integrate the responses. Experiments on mathematical reasoning and coding tasks show that this strategy improves the performance–cost tradeoff compared to existing baselines and can transfer to other advisor model families without additional training.

By Yongjun Kim, Xiaoxiao Li, Jaeho Lee
arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv Computation and Language
Aug 27

Addressing the Reasoning Gap: Mechanistic Circuit-Based Knowledge Editing in Large Language Models

The paper introduces MCircKE, a mechanistic circuit-based knowledge editing framework for large language models. MCircKE identifies the causal circuits involved in a specific reasoning task and surgically updates parameters only within those circuits, thereby addressing the reasoning gap where edited facts are not used in multi-step reasoning. Experiments on the MQuAKE-series benchmarks show that this approach improves multi-hop reasoning performance after knowledge editing.

By Tianyi Zhao, Yinhan He, Wendy Zheng, Chen Chen
arXiv AI
Sep 7

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

The paper surveys efficient reasoning in large language models, contrasting fast intuitive (System 1) and slow deep (System 2) reasoning. It analyzes why System 2 is computationally costly yet more accurate, and why System 1 is efficient but less effective. The survey covers causes of inefficiency, patterns of reasoning behavior, and potential solutions to balance performance and computational budgets, offering actionable insights and an open‑source repository for ongoing research.

By Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong