arXiv Computation and Language

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

The paper investigates time‑incremental continued pretraining (CPT) of large language models using web‑scale data that overlaps across snapshots. Across six open‑weight models, CPT improves factual recall without catastrophic forgetting, while the cost is negligible and data quality outweighs quantity. The study identifies optimal learning rates, shows LoRA can match full CPT, and demonstrates that CPT gains transfer to fine‑tuned models.

arXiv AI
Sep 4

Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning

The paper introduces LRE (Learned Relevance Eviction), a lightweight, CPU‑only, language‑model‑free scorer that learns which parts of an agent’s interaction history are task‑critical and preserves them verbatim. In experiments, LRE matches or surpasses baseline eviction policies on accuracy‑cost trade‑offs, recovers 93% of full‑history accuracy, reduces worst‑case prompt size by 52%, and outperforms dense and token‑pruning encoders in conversational memory while being 295–1569× smaller. The method also achieves superior budgeted answer quality on LoCoMo reading and can be trained annotation‑free, recovering 95% of supervised scorer performance.

By Nusrat Jahan Lia, Aritra Mazumder
arXiv AI
Aug 11

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.

By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
arXiv Computation and Language
Aug 24

Index SLM Technical Report

arXiv:2607.09885v3 Announce Type: replace Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundatio...

By Tianjiao Li, Lusheng Zhang, Shien He, Xiaojing Liu, Tianxing Yan, Mengran Yu, Ziang Cui, Kai Zhao, Xipeng Wang, Yang Liu, Yuxin Li
arXiv AI
Aug 19

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

The paper introduces Spaced Repetition Training (SRT), a continual learning framework that schedules sample rehearsal using the SM-2 algorithm. SRT tracks per-example review states and maps perplexity to recall quality, allowing the training loop to decide which examples to replay and when. Experiments on Wikipedia and code corpora show that SRT improves the stability‑plasticity trade‑off, recovers 5–37 percentage points of lost old‑knowledge accuracy, and preserves benchmark performance better than naive continual pre‑training or uniform replay.

By Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi