Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
arXiv:2511. 20892v4 Announce Type: replace Abstract: Large language models (LLMs) often produce incorrect or outdated content after being employed.
arXiv:2303. 01421v2 Announce Type: replace-cross Abstract: Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks.
arXiv:2511. 20892v4 Announce Type: replace Abstract: Large language models (LLMs) often produce incorrect or outdated content after being employed.
arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.
arXiv:2608. 12419v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications.
The paper introduces the Continuous Memory Machine (CMM), a recurrent neural network that uses separate matrix-valued short‑term and long‑term memory states. Short‑term memory tracks recent neural activity with neuron‑level models for rapid computation, while long‑term memory stores information for later use; both are jointly updated by a Transformer that allows bidirectional read‑write operations. Experiments on algorithmic, in‑context learning, and recurrent reasoning tasks show that CMM outperforms many baselines and generalizes better than previous memory‑augmented networks, while maintaining interpretable attention patterns from the Continuous Thought Machine.
arXiv:2608.17050v3 Announce Type: replace-cross Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access...
arXiv:2609.16540v1 Announce Type: cross Abstract: State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear comput...
arXiv:2606. 15734v1 Announce Type: cross Abstract: Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing catastrophic forgetting and degrading general capabilities.
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
arXiv:2610.08630v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, u...
arXiv:2502.09245v3 Announce Type: replace Abstract: In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standar...
The paper investigates how Engram-style hashed memory can be transferred between different language model backbones. By freezing a memory table trained on a source model and attaching it to a target model with only a lightweight reader, the authors find that both the memory content and correct addressing are important, but the reader must be aligned to the target to make the memory useful. In question‑answering experiments, a dual‑layer, four‑branch reader nearly matches same‑model performance, and when the reader interface is directly compatible, the frozen memory alone provides substantial benefit, with optional reader adaptation offering further gains.