arXiv Machine Learning By Guangyue Peng, Tao Ge, Wen Luo, Wei Li, Houfeng Wang

Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory

Read the original on arXiv Machine Learning →

arXiv:2303. 01421v2 Announce Type: replace-cross Abstract: Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.

By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
arXiv AI
2d ago

Continuous Memory Machines

The paper introduces the Continuous Memory Machine (CMM), a recurrent neural network that uses separate matrix-valued short‑term and long‑term memory states. Short‑term memory tracks recent neural activity with neuron‑level models for rapid computation, while long‑term memory stores information for later use; both are jointly updated by a Transformer that allows bidirectional read‑write operations. Experiments on algorithmic, in‑context learning, and recurrent reasoning tasks show that CMM outperforms many baselines and generalizes better than previous memory‑augmented networks, while maintaining interpretable attention patterns from the Continuous Thought Machine.

By Ciaran Regan, Kai Arulkumaran, Luke Darlow, Stefania Druga, Sebastian Risi, Llion Jones