arXiv Machine Learning

Continual Learning Mechanisms Compose for Long-Horizon Memorization

arXiv AI
Aug 11

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.

By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
arXiv AI
Aug 20

Forgetting, plasticity, and co-observation: a third facet of continual learning

The paper argues that catastrophic forgetting and loss of plasticity alone cannot explain why naive sequential training underperforms offline joint training. It introduces data co-observation as a third factor, showing that observing training data together consistently improves performance across supervised and self-supervised settings. The study also reinterprets common continual learning methods, suggesting that memory replay’s success stems from restoring co-observation benefits rather than merely mitigating forgetting.

By Timm Hess, Abhishek Jha, Gido M. van de Ven, Tinne Tuytelaars