RL Forgets! Towards Continual Policy Optimization
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
arXiv:2607. 02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined.
arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.
arXiv:2607. 27260v1 Announce Type: new Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge.
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
arXiv:2605. 20247v2 Announce Type: replace-cross Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).
arXiv:2607. 15587v1 Announce Type: new Abstract: Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch.
The paper introduces Spaced Repetition Training (SRT), a continual learning framework that schedules sample rehearsal using the SM-2 algorithm. SRT tracks per-example review states and maps perplexity to recall quality, allowing the training loop to decide which examples to replay and when. Experiments on Wikipedia and code corpora show that SRT improves the stability‑plasticity trade‑off, recovers 5–37 percentage points of lost old‑knowledge accuracy, and preserves benchmark performance better than naive continual pre‑training or uniform replay.
The paper introduces Spaced Repetition Training (SRT), a continual learning framework that adapts review scheduling for language models by using the SM-2 algorithm to decide which past examples to replay. SRT tracks per-example review states and maps perplexity to a recall-quality signal, allowing the model to retain old knowledge while consolidating new information without changing the underlying model or training objective. Experiments on Wikipedia and code corpora show that SRT improves the stability-plasticity trade‑off, recovers 5–37 percentage points of lost accuracy, and maintains benchmark performance better than naive continual pre‑training or uniform replay; similar benefits are observed in vision and tabular data when an appropriate recall signal is used.
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised p...
The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.
arXiv:2606. 06032v1 Announce Type: new Abstract: Catastrophic forgetting is commonly interpreted as the irreversible erasure of previously acquired knowledge during sequential learning.
arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.
The paper introduces Continual Reasoning Gym, a continual‑RLVR environment that sequences text and visual reasoning tasks. It finds that while sequential RLVR shows modest forgetting, its final performance lags behind multitask RLVR (MTRL) because forgetting explains only part of the gap. To bridge this, the authors propose Continual Prompt Replay (CPR), which replays previous‑task prompts and regenerates responses with the current policy, achieving on average MTRL‑level performance.