arXiv Machine Learning By Malavika Suresh, Ikechukwu Nkisi-Orji, Nirmalie Wiratunga

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 00630v1 Announce Type: new Abstract: Achieving continual learning (CL) with deep neural networks requires balancing stability and plasticity while enabling knowledge transfer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Parameter Importance-Driven Continual Learning for Foundation Models

The paper introduces PIECE, a Parameter Importance-Driven Continual Learning method that selectively updates only 0.1% of core parameters to preserve general abilities while learning new domain knowledge. PIECE employs two importance estimators—PIECE‑F using Fisher Information and PIECE‑S combining gradient and curvature information—to guide updates. Experiments on three language models and two multimodal models demonstrate that PIECE maintains general capabilities and achieves state‑of‑the‑art continual learning performance without accessing prior training data or adding parameter overhead.

By Lingxiang Wang, Hainan Zhang, Zhiming Zheng
arXiv AI
Jun 16

Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift

arXiv:2606. 15734v1 Announce Type: cross Abstract: Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing catastrophic forgetting and degrading general capabilities.

By Weihang Su, Jiacheng Kang, Jingyan Xu, Qingyao Ai, Jianming Long, Hanwen Zhang, Bangde Du, Xinyuan Cao, Min Zhang, Yiqun Liu
arXiv AI
Aug 11

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.

By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin