arXiv AI

MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller

arXiv Machine Learning
Aug 28

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

The paper investigates whether the high costs of training chain-of-thought reasoning models can be reduced through algorithmic design. It introduces an autocurriculum approach that lets the model select which problems to focus on during training, showing that this method provably improves both supervised fine‑tuning and reinforcement learning. For supervised fine‑tuning, autocurriculum requires exponentially fewer reasoning demonstrations by targeting prompts where the model struggles, while for reinforcement learning it decouples computational cost from the quality of the reference model, making the burn‑in cost nearly independent of target accuracy.

By Nived Rajaraman, Audrey Huang, Miro Dudik, Robert Schapire, Dylan J. Foster, Akshay Krishnamurthy
arXiv AI
3d ago

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

The paper demonstrates that fine‑tuning reasoning models to predict their own confidence at intermediate steps—using only 600 self‑supervised examples—substantially improves inference efficiency. Without adding any explicit stopping or length penalties, the models generate up to 25 % fewer tokens while maintaining accuracy on mathematical, scientific, and coding benchmarks across several architectures. The study finds that confidence supervision preserves the models’ high‑level reasoning structure rather than merely suppressing specific behaviors.

By Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan
arXiv AI
Jun 9

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.

By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su
arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv AI
Sep 18

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

When2Think introduces a post‑training framework that dynamically allocates reasoning depth in Large Reasoning Models based on instance difficulty. The method uses Instance‑level Difficulty‑Aware Control (IDAC) to shape rewards with pre‑computed accuracy and token usage statistics, enabling stable, critic‑free optimization without learned reward models. Experiments on mathematical benchmarks show that When2Think improves accuracy‑efficiency trade‑offs, achieving higher Pass@3 scores while reducing token usage compared to baseline models.

By Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak
arXiv AI
Aug 11

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning

arXiv:2505. 24273v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yielding significant improvements on challenging reasoning tasks through extended chains of thought.

By Hongyi James Cai, Junlin Wang, Xiaoyin Chen, Bhuwan Dhingra
arXiv Computer Vision
Aug 28

Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

The paper introduces Echo-GRPO, a method that rewrites privileged reasoning traces into a model’s own idiolect to align off‑policy supervision with the student policy’s vocabulary. By preserving semantics through Dual‑Reference Decoding, Echo‑GRPO mitigates gradient clipping on critical reasoning tokens and improves reasoning distillation. The approach is instantiated as VideoEcho‑R1 for video reasoning, yielding consistent gains across multiple multimodal LLM backbones and benchmarks, and it can be applied as a plug‑in to both RL and supervised fine‑tuning frameworks.

By Ji Soo Lee, Jinyoung Park, Seohyun Lee, Jongha Kim, Joonmyung Choi, Jinsung Yoon, Hyunwoo J. Kim