arXiv:2606. 17945v1 Announce Type: new Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered.
By Liangkai Hang, Junjie Yao, Zhiyu Li, Feiyu Xiong, Hongkang Yang, Zhi-Qin John Xu
arXiv:2609.38764v1 Announce Type: new
Abstract: Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract,...
By Zachary Shinnick, Hemanth Saratchandran, Damien Teney, Anton van den Hengel
The paper investigates how different forms of compressed chain‑of‑thought (CoT) reasoning—Explicit, Composed, and Implicit—affect large language model (LLM) performance after supervised fine‑tuning (SFT). Using a synthetic compositional reasoning task, the authors show that coarser CoT requires more SFT data, that Composed and Implicit CoT benefit more from data scaling (with Composed also benefiting from repetition), and that reinforcement learning with verifiable rewards (RLVR) can decompose compressed steps learned during SFT. Additionally, unidirectional CoT ordering improves generalization on longer sequential tasks.
By Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
The paper reports that in on‑policy distillation for large language models, reasoning performance can be improved by supervising only a tiny fraction of generated tokens—sometimes just one or two tokens per reasoning trajectory, about 0.05% of all tokens. This sparse supervision consistently matches or exceeds full‑token training across nine teacher‑student setups on mathematical reasoning, and is also validated on coding reasoning, Llama models, and PPO‑based reinforcement learning with verifiable reward. The findings suggest that effective post‑training does not require token‑intensive supervision and may align more closely with natural learning processes that focus on critical reasoning steps.
By Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane
The study evaluates synthetic pre‑pretraining (PPT) across a wide range of models (500 M–7 B parameters) and training budgets (up to 100 B tokens). Results show that PPT consistently improves downstream performance and token efficiency, saving at least 21 B tokens at the 3 B scale, but these gains do not appear to stem from a grammatical prior. Instead, PPT benefits arise from tasks that enhance long‑range retrieval, and the improvements remain robust across diverse data mixtures, diminishing only when web text is omitted.
By Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal \v{S}tef\'anik, Aline Villavicencio, Nikolaos Aletras
arXiv:2606. 25987v1 Announce Type: cross Abstract: Large language models (LLMs) attain remarkable surface fluency on code, yet they neither formally guarantee the syntactic validity of their output nor leverage the hierarchical structure defining the target language.
By Alexandre Bouayad