The paper introduces Spaced Repetition Training (SRT), a continual learning framework that schedules sample rehearsal using the SM-2 algorithm. SRT tracks per-example review states and maps perplexity to recall quality, allowing the training loop to decide which examples to replay and when. Experiments on Wikipedia and code corpora show that SRT improves the stability‑plasticity trade‑off, recovers 5–37 percentage points of lost old‑knowledge accuracy, and preserves benchmark performance better than naive continual pre‑training or uniform replay.
By Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
arXiv:2607. 15587v1 Announce Type: new Abstract: Continual learning studies how deployed language models can continually acquire new tasks without expensive retraining from scratch.
By Yang Meng, Zhenya Liu, Zhuokai Zhao, Yuxin Chen
arXiv:2607. 04969v1 Announce Type: new Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve both model performance and sample efficiency.
By Jingwei Zuo, Cong Zeng, Ilyas Chahed, Maksim Velikanov, Dhia Eddine Rhaiem, Pasquale Balsebre, Abhay Kumar, Younes Belkada, Hakim Hacid
arXiv:2609.40089v1 Announce Type: new
Abstract: Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabiliti...
By Lukas Thede, Shengzhuang Chen, Stefan Winzeck, Matthias Bethge, Zeynep Akata, Jonathan Richard Schwarz
arXiv:2607. 02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined.
By Qianyu Chen, Canran Xiao, Runxuan Tang
arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.
By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin