arXiv Machine Learning

The Dichotomy Between Pattern Recognition and Step-by-Step Reasoning

The paper argues that pattern recognition and step‑by‑step reasoning lie on a spectrum, with large language models (LLMs) learning the latter when the next token depends on only a few preceding tokens. It formalises reasoning traces as paths on a De Bruijn graph, showing that the number of edges is far smaller than the number of possible traces, making step‑by‑step reasoning sample‑efficient. Experiments fine‑tuning Qwen2.5‑1.5B‑Instruct demonstrate that a moderate density of states balances accuracy and robustness, and that real‑world models like Qwen3 retain most of their performance even when attention is limited to a small sliding window.

arXiv AI
Aug 11

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

arXiv:2608. 07885v1 Announce Type: new Abstract: Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain.

By Agamdeep Singh, Srishti Gautam, Priyanshu Gupta, Nikita Mehrotra, Tanmay Bakshi, Sumit Gulwani
arXiv AI
Oct 2

Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models

The paper introduces masked self‑distillation, a post‑training framework that trains a language model to internalize portions of its own intermediate reasoning traces. By varying the fraction of trace internalized, the authors demonstrate that models can achieve higher inference efficiency and improved task performance on math and graph‑coloring problems. Experiments on Qwen3‑4B and Qwen3‑8B show that the method generalizes well to in‑domain out‑of‑distribution cases without catastrophic forgetting, and that supervised fine‑tuning alone can reduce trace length at the expense of generalization.

By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv Machine Learning
Jun 29

Learning to Reason with Curriculum II: Compositional Generalization

arXiv:2606. 27721v1 Announce Type: new Abstract: Compositional generalization, the ability to solve complex problems by combining solutions to simpler sub-problems, is a fundamental capability of both natural and artificial intelligence, and a key mechanism underlying chain-of-thought reasoning.

By Nived Rajaraman, Audrey Huang, Miroslav Dudik, Robert Schapire, Dylan Foster, Akshay Krishnamurthy
arXiv AI
Jun 15

Fractured Chain-of-Thought Reasoning

arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.

By Baohao Liao, Hanze Dong, Yuhui Xu, Doyen Sahoo, Christof Monz, Junnan Li, Caiming Xiong