Small Initialization Matters for Large Language Models
arXiv:2606. 17945v1 Announce Type: new Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered.
arXiv:2608. 04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models.
arXiv:2606. 17945v1 Announce Type: new Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered.
arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
arXiv:2602. 11852v2 Announce Type: replace Abstract: While state-of-the-art language models (LMs) surpass most humans in certain domains, their reasoning remains largely opaque, reducing trust and increasing the risk of deception and hallucination.
arXiv:2601. 22588v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaque, and sensitive to prompt design.
arXiv:2605. 30580v2 Announce Type: replace-cross Abstract: Speculative decoding is a popular technique for large language model (LLM) inference, enabling faster generation by drafting multiple tokens with a smaller draft model.
The paper investigates whether large language models (LLMs) follow Occam's Razor when performing inductive and abductive reasoning. It introduces a synthetic framework for generating questions that require both types of reasoning and a new automated metric to evaluate the simplicity and correctness of generated hypotheses. Experiments show that while LLMs can handle simple scenarios, they struggle with complex world models and producing high‑quality, simplest hypotheses, even when using advanced reasoning techniques.
arXiv:2603. 03031v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved strong complex reasoning capabilities through Chain-of-Thought (CoT) reasoning.
The paper demonstrates that frontier language models can be prompted to expose their internal chain-of-thought reasoning via a simple custom tool. By comparing these extracted traces to native reasoning on open-source models, the authors confirm that the externalized reasoning aligns with genuine reasoning and outperforms no-reasoning baselines across math, science, and code tasks. They further analyze the structure of the reasoning, noting token-efficient, directed reasoning in models like GPT‑6 Astra, which externalizes only crucial steps while handling elementary ones internally.
arXiv:2609.23367v1 Announce Type: cross Abstract: FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising...
arXiv:2607. 18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable.
The paper investigates how transformers acquire deep semantic dependencies, proposing a mechanistic framework that frames learning as a competition between surface statistics and deep semantics. It identifies a "Gradient Starvation" effect that suppresses error signals for sparse semantic dependencies early in training, delaying structural reasoning until a sudden phase transition. The study also explains the success of Chain-of-Thought strategies and introduces a topology‑aligned contrastive objective that improves variable binding performance by more than twice the gain of standard fine‑tuning.
arXiv:2606. 31779v1 Announce Type: new Abstract: Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token.