arXiv Machine Learning By Jingwei Zuo, Cong Zeng, Ilyas Chahed, Maksim Velikanov, Dhia Eddine Rhaiem, Pasquale Balsebre, Abhay Kumar, Younes Belkada, Hakim Hacid

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

Read the original on arXiv Machine Learning →

arXiv:2607. 04969v1 Announce Type: new Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited high-quality data can improve both model performance and sample efficiency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 18

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

The paper introduces Spaced Repetition Training (SRT), a continual learning framework that adapts review scheduling for language models by using the SM-2 algorithm to decide which past examples to replay. SRT tracks per-example review states and maps perplexity to a recall-quality signal, allowing the model to retain old knowledge while consolidating new information without changing the underlying model or training objective. Experiments on Wikipedia and code corpora show that SRT improves the stability-plasticity trade‑off, recovers 5–37 percentage points of lost accuracy, and maintains benchmark performance better than naive continual pre‑training or uniform replay; similar benefits are observed in vision and tabular data when an appropriate recall signal is used.

arXiv AI
Aug 19

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

The paper introduces Spaced Repetition Training (SRT), a continual learning framework that schedules sample rehearsal using the SM-2 algorithm. SRT tracks per-example review states and maps perplexity to recall quality, allowing the training loop to decide which examples to replay and when. Experiments on Wikipedia and code corpora show that SRT improves the stability‑plasticity trade‑off, recovers 5–37 percentage points of lost old‑knowledge accuracy, and preserves benchmark performance better than naive continual pre‑training or uniform replay.

By Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
arXiv AI
Aug 11

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.

By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
arXiv AI
Aug 20

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

The paper investigates training-time data augmentation as a regularizer for autoregressive language model pretraining in data‑constrained, compute‑abundant settings. It introduces three orthogonal augmentation categories—token‑level noise, sequence permutations, and target offset prediction—and shows through systematic ablations that each category delays overfitting and reduces validation loss, with random token replacement performing best individually. Combining augmentation categories further lowers the minimum validation loss, demonstrating that such augmentations mitigate data inefficiency in autoregressive pretraining.

By Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang