arXiv Machine Learning
4d ago

Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

The paper introduces a Fisher-guided submodular data selection method for continual pre‑training of large language models, addressing the forgetting problem that arises when a target‑domain corpus overwrites pretrained knowledge. By decomposing candidate gradients into anchor and frontier components and optimizing a log‑determinant submodular objective, the method selects data that both improves target‑domain performance and limits forgetting. Experiments on TinyLlama‑1.1B and Llama‑3.1‑8B show that the selector achieves superior adaptation and forgetting control while using ten times fewer tokens than traditional replay strategies.

By Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan
arXiv AI
Sep 28

Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models

The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.

By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama