arXiv AI
Sep 7

Extremely Sparse Supervision Incentivizes Reasoning Ability

The paper reports that in on‑policy distillation for large language models, reasoning performance can be improved by supervising only a tiny fraction of generated tokens—sometimes just one or two tokens per reasoning trajectory, about 0.05% of all tokens. This sparse supervision consistently matches or exceeds full‑token training across nine teacher‑student setups on mathematical reasoning, and is also validated on coding reasoning, Llama models, and PPO‑based reinforcement learning with verifiable reward. The findings suggest that effective post‑training does not require token‑intensive supervision and may align more closely with natural learning processes that focus on critical reasoning steps.

By Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane
arXiv AI
Jul 28

DeepLook: Deeper Thinking with Lookahead

arXiv:2607. 22602v1 Announce Type: new Abstract: Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter scaling alone.

By Tingxin Yang, Zefeng Wang, Mengyue Wang, Xingcheng Zhou, Yunpu Ma
arXiv Computation and Language
Sep 1

Unlocking Fine-Grained Translation Quality Estimation in LRMs through Mutually Boosting Implicit and Explicit Reasoning

arXiv:2605.31378v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) still struggle with fine-grained translation quality estimation (QE), even with long reasoning chains. We argue that...

By Renfei Dang, Xinye Wang, Zhejian Lai, Weilu Xu, Shimin Tao, Daimeng Wei, Min Zhang, Shujian Huang
arXiv AI
Sep 11

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

The paper introduces Trimmed Logit-Gap SFT (TrimSFT), a token-level reweighting strategy that adjusts supervised fine-tuning loss based on the logit gap between the correct token and its strongest competitor. TrimSFT trims supervision from tokens that are either already mastered (large logit gap) or poorly supported (small or negative logit gap), focusing learning on tokens with intermediate logit gaps. Experiments on six base models across five mathematical reasoning benchmarks show that TrimSFT consistently outperforms standard SFT, achieving the best average performance on five of six models and up to +26.9 points on MATH500.

By Yaning Jia, Chunhui Zhang, Wenxuan Xu, Xingjian Diao, Xiaoyuan Wang, Soroush Vosoughi