Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
arXiv:2608. 12476v1 Announce Type: new Abstract: Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim.
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it emits that stale value as a confident answer; give the same model an empty memory and it abstains.
The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.
arXiv:2609. 04875v1 Announce Type: cross Abstract: Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache.
arXiv:2608. 07429v1 Announce Type: new Abstract: Long-term memory enables language agents to reuse past facts, preferences, and task experience.
The paper investigates how persistent memory in AI agents can lead to over‑trust in stale facts, creating a "Memory Trust Gap" that worsens as model capability increases. Using a benchmark with Benefit and Safety suites across Qwen3 models of varying sizes, the authors show that larger models are more prone to harmful over‑trust, especially when metadata is absent or misleading. They also demonstrate that mitigation strategies such as exposing metadata or pre‑resolving conflicts improve accuracy, but the effectiveness depends on model size and dataset.
arXiv:2606. 06240v1 Announce Type: cross Abstract: Persistent memory for an LLM agent is a write-heavy substrate: every belief update is a versioned write, and a new claim may contradict a stored one.
arXiv:2609.22243v1 Announce Type: new Abstract: Input-conditioned neural interventions raise a runtime question: what persists when one behavioral specification admits multiple actions whose validity...
The paper investigates how agents that inherit consolidated memory can mistakenly rely on constraints that were once true but have since been superseded by newer records. By modeling supersession explicitly and limiting verification to two records, the study shows that native allocation often leads to stale-consistent decisions, while reallocating one verification slot to the critical provenance path markedly improves consistency. The results suggest that memory systems may need separate freshness or supersession signals beyond relevance to avoid such errors.
arXiv:2608. 05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression.
arXiv:2609.08258v1 Announce Type: new Abstract: Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history through soft revocation: a contradicted fact...
arXiv:2607. 23806v1 Announce Type: cross Abstract: Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output.