arXiv:2607. 17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks.
By Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen
The paper reports that in on‑policy distillation for large language models, reasoning performance can be improved by supervising only a tiny fraction of generated tokens—sometimes just one or two tokens per reasoning trajectory, about 0.05% of all tokens. This sparse supervision consistently matches or exceeds full‑token training across nine teacher‑student setups on mathematical reasoning, and is also validated on coding reasoning, Llama models, and PPO‑based reinforcement learning with verifiable reward. The findings suggest that effective post‑training does not require token‑intensive supervision and may align more closely with natural learning processes that focus on critical reasoning steps.
By Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane
The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali.
arXiv:2607. 22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time.
By Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
The paper argues that tokenization in language models should be viewed as an output supervision decision rather than merely input preprocessing. In autoregressive models, the granularity of the tokenizer determines the supervision signal the model receives, influencing learning difficulty, internal representations, and task performance. Experiments on numeric reasoning show that output tokenization, rather than input tokenization, drives differences in performance and training dynamics, and a survey of recent CL papers reveals that tokenization choices are rarely reported or acknowledged.
By Tanja Baeumel, Josef van Genabith, Simon Ostermann
arXiv:2608.28600v1 Announce Type: new
Abstract: Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their...
By Jonghyun Song, Sangjun Song, Minjae Oh, Haesung Pyun, Sungsik Lee, Yohan Jo
The paper introduces RATIO, a framework for improving quantized reasoning models by identifying overthinking tokens and applying token-specific penalties. It uses Quantization-aware Reasoning Behavior Analysis to detect problematic tokens and Token-Specific Penalty Determination to assign penalties without extra training. Experiments show RATIO outperforms existing methods, boosting accuracy by up to 9.8 points and shortening chain-of-thought length by up to 51.3%.
By Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang
arXiv:2603.02504v3 Announce Type: replace
Abstract: Large Language Models (LLMs) achieve strong performance on natural language tasks but remain unreliable in mathematical reasoning, frequently gener...
By Pratibha Zunjare, Michael Hsiao
arXiv:2609.37066v1 Announce Type: cross
Abstract: Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed...
By Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu, Said Mammar, Pascal Bouvry
The paper introduces Abstract Token Curriculum (ATC), a curriculum learning framework that enables large language models to develop continuous intermediate representations—referred to as abstract thoughts—without explicit supervision or manual scratchpad design. ATC incrementally raises problem difficulty through a sequence of distributions, guiding models to focus attention on the most informative tokens for predicting subsequent tokens. The authors provide theoretical analysis for parity function learning with single‑layer softmax attention and demonstrate ATC’s effectiveness on graph reachability and arithmetic learning tasks.
By Khashayar Gatmiry, Avrajit Ghosh, Parsa Mirtaheri, Jason D. Lee, Nika Haghtalab, Emmanuel Abbe, Peter Bartlett
The paper investigates on‑policy self‑distillation (OPSD) as a method to enhance reasoning in language models, focusing on mathematical reasoning across models from 0.6B to 8B parameters. Through controlled experiments and token‑level analysis, the authors find that OPSD’s effectiveness depends on alignment between the teacher’s reasoning mode and the full teacher prefix, rather than on privileged semantics alone. They observe that OPSD only improves reasoning in limited compatibility regimes, while often causing length growth, degradation, or behavioral collapse, and that the teacher’s signal is unstable and not predictive of downstream performance.
By Yang Li, Gongle Xue, Yuheng Yuan, Yijia Guo, Shizhe Zhang, Liwen Hu, Lei Ma
arXiv:2609.25438v1 Announce Type: new
Abstract: Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine...
By Henry Kvinge