arXiv AI By Zi Wang, Xingqiao Wang, Emmanuel Addai, Devika Ambekar, Xiaowei Xu

The Right Memory in the Wrong Context: Verifying Retrieval Admissibility in Long-Term Agent Memory

Read the original on arXiv AI →

The paper presents a retrieval‑admissibility verification framework for long‑term memory agents, classifying each memory‑query pair as admissible, inadmissible, or unresolved. It evaluates the framework on public benchmarks (RHELM and MemOps), showing improved anchor recall and reduced exact similarity errors, while also revealing that existing verifiers miss certain inadmissible exposures. The study highlights the need for separate checks on candidate support, admissibility, prompt exposure, and answer disclosure to ensure safe memory retrieval.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

The paper introduces Causal Memory Policy (CMP), a framework that identifies the utility of memories in memory‑augmented language models by intervening on retrieval rather than on storage. CMP reserves fixed context slots for memories sampled with known propensities and estimates utility using self‑normalized inverse propensity weighting, providing unbiased estimates and exact variance. Experiments show that CMP improves discrimination between required and non‑required memories and reveals that identified utility alone is insufficient for retention decisions across unseen queries.

By Arman Behnam, Binghui Wang
arXiv AI
Sep 11

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall (FR) introduces an ontology-driven policy layer that categorizes personal facts into over ten behavioral types and applies tailored lifecycle rules—such as differential decay, supersession, and event-time validity—to manage memory persistence in large language models. The FR-Bank implementation, independent of underlying infrastructure, achieves a 76.9% pass rate on the new LifecycleBench benchmark and improves LongMemEval-S performance, while significantly reducing confabulation rates compared to prior systems. Ablation studies show that the generic lifecycle metadata drives correctness, whereas the behavioral ontology enhances calibration and reduces downstream hallucinations.

By Ansuman Mullick, Eray T\"uz\"un
arXiv AI
Jun 16

Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations

arXiv:2606. 15903v1 Announce Type: cross Abstract: Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the control plane that mutates them via supersede, release, purge (largely untested) -- shapes which forgetting failure modes the system recovers.

By Dongxu Yang