arXiv AI
6d ago

Masked Self-Distillation: Internalizing the Chain-of-Thought in Language Models

The paper introduces masked self‑distillation, a post‑training framework that trains a language model to internalize portions of its own intermediate reasoning traces. By varying the fraction of trace internalized, the authors demonstrate that models can achieve higher inference efficiency and improved task performance on math and graph‑coloring problems. Experiments on Qwen3‑4B and Qwen3‑8B show that the method generalizes well to in‑domain out‑of‑distribution cases without catastrophic forgetting, and that supervised fine‑tuning alone can reduce trace length at the expense of generalization.

By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv AI
Aug 11

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

arXiv:2608. 08176v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation.

By Yongkang Yang, Zhezheng Hao, Hong Zhang, Yi Liu, Xiankun Lin, Wence Ji, Fanjunduo Wei, Jiarui Yu, Qiang Lin, Xiaoyun Liang, Hande Dong
arXiv AI
Sep 7

Extremely Sparse Supervision Incentivizes Reasoning Ability

The paper reports that in on‑policy distillation for large language models, reasoning performance can be improved by supervising only a tiny fraction of generated tokens—sometimes just one or two tokens per reasoning trajectory, about 0.05% of all tokens. This sparse supervision consistently matches or exceeds full‑token training across nine teacher‑student setups on mathematical reasoning, and is also validated on coding reasoning, Llama models, and PPO‑based reinforcement learning with verifiable reward. The findings suggest that effective post‑training does not require token‑intensive supervision and may align more closely with natural learning processes that focus on critical reasoning steps.

By Zhishuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu, Pan Xu, Karim Bouyarmane