Hugging Face Trending Papers

Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

The paper demonstrates that off‑policy merging, called grafting, outperforms on‑policy self‑distillation for continual learning. Grafting learns updates on an earlier donor checkpoint, scales the update, and optionally masks sensitive directions, thereby reducing interference with existing capabilities. Across various continual learning scenarios, grafting achieves better new‑task and old‑task performance without costly on‑policy sampling.

Hugging Face Trending Papers
Jul 21

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in both computational infrastructure and cost, especially when considering RL as merely a one-off learning stage.

arXiv AI
Oct 2

Finetuning with Sampling: SFT Learns Better Than You Think

The paper introduces a Markov chain Monte Carlo (MCMC) sampling algorithm that transforms off‑policy traces into more on‑policy data for supervised finetuning (SFT). By tailoring the data distribution rather than changing the learning objective, the method allows SFT to match or surpass traditional post‑training techniques such as reinforcement learning across tasks like scientific skill acquisition, mathematical reasoning, and open‑ended expertise. The resulting models generalize better, forget less, and exhibit strong distributional performance, demonstrating that sampling can serve as a model‑native operator for broader post‑training applications.

By Aayush Karan, Sitan Chen, Yilun Du
arXiv AI
Aug 20

Forgetting, plasticity, and co-observation: a third facet of continual learning

The paper argues that catastrophic forgetting and loss of plasticity alone cannot explain why naive sequential training underperforms offline joint training. It introduces data co-observation as a third factor, showing that observing training data together consistently improves performance across supervised and self-supervised settings. The study also reinterprets common continual learning methods, suggesting that memory replay’s success stems from restoring co-observation benefits rather than merely mitigating forgetting.

By Timm Hess, Abhishek Jha, Gido M. van de Ven, Tinne Tuytelaars
arXiv Machine Learning
6d ago

Learning from the Near Future: Temporal Self-Distillation for RLVR

The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.

By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
arXiv Computation and Language
Sep 28

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv Machine Learning
Sep 15

Realistic Continual Learning Approach using Pre-trained Models

arXiv:2404.07729v2 Announce Type: replace Abstract: Continual learning (CL) evaluates adaptability in learning solutions to retain knowledge. Our research addresses the challenge of catastrophic forg...

By Nadia Nasri, Carlos Guti\'errez-\'Alvarez, Sergio Lafuente-Arroyo, Saturnino Maldonado-Basc\'on, Roberto J. L\'opez-Sastre