arXiv:2511. 20892v4 Announce Type: replace Abstract: Large language models (LLMs) often produce incorrect or outdated content after being employed.
By Xuyuan Liu, Shengyu Chen, Xinshuai Dong, Yanchi Liu, Xujiang Zhao, Haoyu Wang, Yujun Yan, Haifeng Chen, Zhengzhang Chen
arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.
By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin
arXiv:2608. 12419v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications.
By Qiuwu Chen, Zimo Liu, Yuchen Li, Ying Sun, Yifan Zhang, Zhijie Qiu, Zeng You, Ryan Dong, Simeng Ma, Yaofo Chen, Mingkui Tan
The paper introduces the Continuous Memory Machine (CMM), a recurrent neural network that uses separate matrix-valued short‑term and long‑term memory states. Short‑term memory tracks recent neural activity with neuron‑level models for rapid computation, while long‑term memory stores information for later use; both are jointly updated by a Transformer that allows bidirectional read‑write operations. Experiments on algorithmic, in‑context learning, and recurrent reasoning tasks show that CMM outperforms many baselines and generalizes better than previous memory‑augmented networks, while maintaining interpretable attention patterns from the Continuous Thought Machine.
By Ciaran Regan, Kai Arulkumaran, Luke Darlow, Stefania Druga, Sebastian Risi, Llion Jones
arXiv:2608.17050v3 Announce Type: replace-cross
Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access...
By Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
arXiv:2609.16540v1 Announce Type: cross
Abstract: State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear comput...
By William L. Tong, Aryo Lotfi, Emmanuel Abbe, Kostas Vaggelakos, Vishnu Banna, Etai Littwin, Josh Susskind, Cengiz Pehlevan, Eran Malach