RL Forgets! Towards Continual Policy Optimization
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
arXiv:2510. 18874v3 Announce Type: replace Abstract: Adapting language models (LMs) to new tasks via post-training carries the risk of degrading existing capabilities -- a phenomenon classically known as catastrophic forgetting.
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
arXiv:2608. 01743v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central paradigm for large language model (LLM) post-training, but optimization toward new objectives can degrade capabilities already present in the base model.
arXiv:2409. 02228v2 Announce Type: replace Abstract: When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change?
arXiv:2609.06986v1 Announce Type: new Abstract: Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we i...
arXiv:2601.21702v4 Announce Type: replace Abstract: We consider Representation Misdirection (RM), a class of large language model (LLM) unlearning methods that achieve forgetting by redirecting the l...
arXiv:2608. 08673v1 Announce Type: new Abstract: This work explores the relationship between task similarity and catastrophic forgetting in reinforcement learning.
arXiv:2607. 15587v1 Announce Type: new Abstract: Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch.
arXiv:2606. 06320v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities.
arXiv:2511. 04666v4 Announce Type: replace Abstract: A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge as they adapt to new data.