Hugging Face Trending Papers

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution. However, because this supervision is generated via uncontrolled sampling, it provides no diagnostic insight into the model's specific errors or corrective guidance for its individual failure patterns.

arXiv Machine Learning
Jun 18

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

arXiv:2606. 18844v1 Announce Type: new Abstract: Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution.

By Zhilin Huang, Hang Gao, Ziqiang Dong, Yuan Chen, Yifeng Luo, Chujun Qin, Jingyi Wang, Yang Yang, Guanjun Jiang
arXiv AI
Jun 24

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

arXiv:2606. 24064v1 Announce Type: new Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason.

By Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang
arXiv AI
Jun 30

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

arXiv:2606. 30345v1 Announce Type: cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning tasks.

By Haisen Luo, Yiwei Liu, Haoning Wang, Dan Liu, Junxi Yin, Haotian Wang, Lei Zhang, Xiaoyu Tian, Shuaiting Chen, Yuansheng Song, Baoyan Guo, Xiongfei Yan, Bolan Yang, Chengwei Liu, Ming Cui, Jiong Chen
arXiv Computation and Language
Sep 18

Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

Reflective Recovery is a self‑supervised method that turns failed reasoning attempts into training data, enabling large language models to learn how to correct mistakes during inference. By extracting initial segments of erroneous trajectories and using them as prompts, the approach teaches models to recognize and recover from errors without external critics. Experiments show significant accuracy gains on benchmarks such as AIME 2025 and Minerva, and the method overcomes the scaling collapse problem, fostering emergent self‑correction behaviors.

By Qirui Chen, Renjie Pi, Jiahui Gao, Lingpeng Kong
arXiv AI
Jun 2

SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting

arXiv:2604. 10688v2 Announce Type: replace-cross Abstract: On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewards make token-level credit assignment notoriously difficult.

By Binbin Zheng, Xing Ma, Yiheng Liang, Jingqing Ruan, Xiaoliang Fu, Kepeng Lin, Benchang Zhu, Ke Zeng, Xunliang Cai
arXiv Machine Learning
Sep 25

CataOPD: Catalytic On-Policy Distillation for Large Language Model Reasoning

CataOPD introduces a new framework for improving large language model reasoning by combining reinforcement learning and on‑policy distillation. The method treats the teacher as a catalyst that expands the student’s reachability, using Self‑Rescue Routing to find correct trajectories through additional on‑policy sampling and Catalytic‑Guided Self‑Resolution to elicit verified student trajectories. Barrier‑Weighted Internalization further focuses updates on decisive tokens, leading to better performance on unseen problems and improved out‑of‑distribution generalization.

By Wenjin Liu, Chenxi Wang, Jiapu Wang, Zhe Cui, Anh Tuan Luu, Haoran Luo