arXiv Machine Learning

Replay on Demand: An Emergent Curriculum for Balancing Adaptation and Forgetting in Continued Pretraining

Hugging Face Trending Papers
Aug 18

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

The paper introduces Spaced Repetition Training (SRT), a continual learning framework that adapts review scheduling for language models by using the SM-2 algorithm to decide which past examples to replay. SRT tracks per-example review states and maps perplexity to a recall-quality signal, allowing the model to retain old knowledge while consolidating new information without changing the underlying model or training objective. Experiments on Wikipedia and code corpora show that SRT improves the stability-plasticity trade‑off, recovers 5–37 percentage points of lost accuracy, and maintains benchmark performance better than naive continual pre‑training or uniform replay; similar benefits are observed in vision and tabular data when an appropriate recall signal is used.

arXiv AI
Aug 19

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

The paper introduces Spaced Repetition Training (SRT), a continual learning framework that schedules sample rehearsal using the SM-2 algorithm. SRT tracks per-example review states and maps perplexity to recall quality, allowing the training loop to decide which examples to replay and when. Experiments on Wikipedia and code corpora show that SRT improves the stability‑plasticity trade‑off, recovers 5–37 percentage points of lost old‑knowledge accuracy, and preserves benchmark performance better than naive continual pre‑training or uniform replay.

By Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
arXiv Machine Learning
Jul 7

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

arXiv:2607. 04969v1 Announce Type: new Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve both model performance and sample efficiency.

By Jingwei Zuo, Cong Zeng, Ilyas Chahed, Maksim Velikanov, Dhia Eddine Rhaiem, Pasquale Balsebre, Abhay Kumar, Younes Belkada, Hakim Hacid