Hugging Face Trending Papers

Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning

Read the original on Hugging Face Trending Papers →

The paper demonstrates that off‑policy merging, called grafting, outperforms on‑policy self‑distillation for continual learning. Grafting learns updates on an earlier donor checkpoint, scales the update, and optionally masks sensitive directions, thereby reducing interference with existing capabilities. Across various continual learning scenarios, grafting achieves better new‑task and old‑task performance without costly on‑policy sampling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

Hugging Face Trending Papers
Jul 21

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in both computational infrastructure and cost, especially when considering RL as merely a one-off learning stage.