Controlling Reasoning Effort in LLMs
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
Related stories
Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
The paper investigates whether large language models (LLMs) possess coherent, human-like knowledge structures in mathematical reasoning by applying Knowledge Space Theory (KST). Using a KST-based framework, the authors evaluate eight open- and closed-source LLMs and find that they frequently violate knowledge dependencies, fail to leverage related context, and exhibit low overlap in knowledge distributions compared to real human learners. These structural deficiencies remain largely invisible to standard accuracy or LLM-as-judge evaluations, suggesting that current LLMs do not follow a human-like knowledge structure.
Categories of Inference-Time Scaling for Improved LLM Reasoning
And an Overview of Recent Inference-Scaling Papers
Boosting LLM Reasoning via Human-Inspired Reward Shaping
The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
arXiv:2410.23912v3 Announce Type: replace-cross Abstract: The reasoning abilities of large language models (LLMs) have improved with chain-of-thought (CoT) prompting, allowing models to solve complex...
LeAct: Learning to Reason from Expert Actions
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs
Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability. While recent studies have probed internal representations of model solvability beliefs, verbalization has primarily been studied behaviorally rather than as an internal representation, limiting its analysis and manipulation.
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
The paper investigates why self‑distillation can sometimes worsen the reasoning abilities of large language models (LLMs). It finds that the process suppresses the model’s epistemic verbalization—its expression of uncertainty—leading to shorter but less accurate responses in mathematical reasoning tasks. Experiments on several LLMs show performance drops of up to 40%, especially on out‑of‑distribution problems where uncertainty expression is beneficial.
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
arXiv:2608. 03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities.
Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
arXiv:2608. 11829v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning.
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
arXiv:2605. 24396v2 Announce Type: replace Abstract: Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute.
Knowledge Knows, Verbalization Tells: Disentangling Latent Directions for Mathematical Solvability in LLMs
arXiv:2607. 05013v1 Announce Type: cross Abstract: Although LLMs have made significant progress in mathematical reasoning, determining whether a mathematical problem is solvable remains a fundamental yet challenging capability.
