Length Penalties Make Chain-of-Thought Less Monitorable
arXiv:2607. 09786v1 Announce Type: new Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer.
The paper investigates how training large language models to use fewer tokens in chain-of-thought (CoT) reasoning impacts the faithfulness and monitorability of the generated explanations. Three efficiency methods—fixed generation budget, per-example length target, and group-relative length reward—were applied during fine‑tuning, and the resulting models were evaluated on how well their CoT reflects decision processes and whether it signals changes due to input interventions. Results show that while faithfulness generally decreases because models become less consistent, monitorability remains relatively robust, with models still indicating the influence of input changes even when CoT is shortened.
arXiv:2607. 09786v1 Announce Type: new Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer.
arXiv:2602. 20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output.
arXiv:2608. 04771v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost.
arXiv:2605. 24396v2 Announce Type: replace Abstract: Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute.
Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps cause severe overthinking that inflates inference cost. KV-cache compression is a common solution, yet existing reasoning-oriented methods apply a uniform policy across the trajectory and judge compression only by what it removes from the cache.
The paper introduces CARE, a contrastive accuracy reward estimation method that adaptively adjusts reasoning length for large language models. By comparing beneficial length adjustments from online sampled responses, CARE applies adaptive length rewards within Group Relative Policy Optimization without extra hyperparameters or inference cost. Experiments on multiple reasoning benchmarks show that CARE improves Pass@1 by up to 4% while reducing reasoning length by 37%, achieving higher token efficiency.
arXiv:2508. 02178v3 Announce Type: replace Abstract: Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process.
The paper investigates how different forms of compressed chain‑of‑thought (CoT) reasoning—Explicit, Composed, and Implicit—affect large language model (LLM) performance after supervised fine‑tuning (SFT). Using a synthetic compositional reasoning task, the authors show that coarser CoT requires more SFT data, that Composed and Implicit CoT benefit more from data scaling (with Composed also benefiting from repetition), and that reinforcement learning with verifiable rewards (RLVR) can decompose compressed steps learned during SFT. Additionally, unidirectional CoT ordering improves generalization on longer sequential tasks.
arXiv:2608. 03550v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities.
arXiv:2606. 26935v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning.
arXiv:2603. 01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning.
The study investigates whether large language models (LLMs) that allocate extra computation during inference—termed reasoning models—reduce classic decision biases compared to their non‑reasoning counterparts. Using 30 vignettes covering six cognitive biases and varying token budgets up to 8,192 tokens, the authors find that reasoning models are not less biased, and increased deliberation does not reliably diminish bias magnitude. Only anchoring showed a bias in the human direction, while other biases either remained unchanged or moved further from human patterns, suggesting that test‑time reasoning does not guarantee rationality.