arXiv Machine Learning

RL Forgets! Towards Continual Policy Optimization

arXiv:2607. 04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks.

arXiv Machine Learning
Jun 2

Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning

arXiv:2603. 11653v2 Announce Type: replace Abstract: Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that can adapt in openended, evolving environments.

By Jiaheng Hu, Jay Shim, Chen Tang, Yoonchang Sung, Bo Liu, Peter Stone, Roberto Martin-Martin
arXiv AI
Jul 3

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

arXiv:2607. 02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined.

By Qianyu Chen, Canran Xiao, Runxuan Tang
arXiv Machine Learning
Sep 25

Beyond Forgetting: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

The paper introduces Continual Reasoning Gym, a continual‑RLVR environment that sequences text and visual reasoning tasks. It finds that while sequential RLVR shows modest forgetting, its final performance lags behind multitask RLVR (MTRL) because forgetting explains only part of the gap. To bridge this, the authors propose Continual Prompt Replay (CPR), which replays previous‑task prompts and regenerates responses with the current policy, achieving on average MTRL‑level performance.

By Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan
arXiv Machine Learning
Aug 20

Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

The paper introduces Continual Reasoning Gym, a continual reinforcement learning with verifiable rewards (RLVR) environment that sequences text and visual reasoning tasks. It finds that sequential RLVR suffers modest forgetting and underperforms multitask RLVR, but that shared reasoning structures can be leveraged. The authors propose Continual Prompt Replay (CPR), which replays previous-task prompts and regenerates responses, achieving performance comparable to multitask RLVR.

By Lirui Luo, Guoxi Zhang, Hongming Xu, Rongqing Li, Cong Fang, Lifeng Fan
arXiv AI
Jun 19

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models

arXiv:2510. 21978v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for contemporary language and vision-language models.

By Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei