Training Continuous Chain of Thought Models: A Tale of Two Regimes
arXiv:2607. 16972v1 Announce Type: new Abstract: Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations.
The paper introduces Abstract Token Curriculum (ATC), a curriculum learning framework that enables large language models to develop continuous intermediate representations—referred to as abstract thoughts—without explicit supervision or manual scratchpad design. ATC incrementally raises problem difficulty through a sequence of distributions, guiding models to focus attention on the most informative tokens for predicting subsequent tokens. The authors provide theoretical analysis for parity function learning with single‑layer softmax attention and demonstrate ATC’s effectiveness on graph reachability and arithmetic learning tasks.
arXiv:2607. 16972v1 Announce Type: new Abstract: Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations.
arXiv:2412.06769v4 Announce Type: replace Abstract: Large language models (LLMs) are typically constrained to reason in the language space, where they express the reasoning process through a chain-of...
arXiv:2509. 04027v3 Announce Type: replace Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs).
arXiv:2505.22635v2 Announce Type: replace-cross Abstract: A common approach for teaching large language models (LLMs) to reason is to train on chain-of-thought (CoT) traces of in-distribution reasoni...
arXiv:2608. 09432v1 Announce Type: cross Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order.
arXiv:2606. 13862v1 Announce Type: cross Abstract: Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation.
arXiv:2609.25438v1 Announce Type: new Abstract: Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine...
arXiv:2607. 19358v1 Announce Type: new Abstract: Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm.
arXiv:2608.31069v1 Announce Type: new Abstract: Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and fo...
arXiv:2607. 10386v1 Announce Type: cross Abstract: Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient.
The paper investigates how different forms of compressed chain‑of‑thought (CoT) reasoning—Explicit, Composed, and Implicit—affect large language model (LLM) performance after supervised fine‑tuning (SFT). Using a synthetic compositional reasoning task, the authors show that coarser CoT requires more SFT data, that Composed and Implicit CoT benefit more from data scaling (with Composed also benefiting from repetition), and that reinforcement learning with verifiable rewards (RLVR) can decompose compressed steps learned during SFT. Additionally, unidirectional CoT ordering improves generalization on longer sequential tasks.
arXiv:2602. 03542v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are often specified in natural language.