arXiv AI

Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents

The paper introduces the Distributed‑Evidence Paradox, where long‑running LLM agents compress past interactions into persistent memories that may not be fully supported by the interaction history. It defines three key requirements—evidence scope, compositional validity, and admission reliability—and proposes DerivAudit, a framework that checks whether a memory is truly supported by the available history. Experiments on two memory corpora show that expanding the evidence base can recover support for many memories, yet many remain unsupported, and broader evidence alone does not guarantee reliable admission.

arXiv Computation and Language
Sep 1

Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents

Agent Zero Memory is a provenance‑aware long‑term memory system for large language model agents that distills user interactions into three parallel memory structures: an episodic timeline, an associative entity‑event knowledge graph, and a semantic, citation‑locked hierarchical documentary memory. Retrieval is performed via an intent gate, source router, and concurrent searches across the three systems, producing integrated, cited answers that exclude fabrication and require evidence the reader has opened. The system achieves state‑of‑the‑art performance on LongMemEval (95.60%) and LoCoMo (93.60%) while offering a favorable accuracy‑cost‑latency trade‑off across multiple backbone LLMs.

By Ming Wu, Pengyuan Zhu
arXiv Computation and Language
4d ago

Memory Consolidation Flattens the Temporal Shape of User Facts

The paper introduces the concept of aspectual flattening, where long‑term memory systems convert conversational statements into concise notes that omit temporal cues. Using the LAPSE benchmark, the authors show that several memory writer models consistently flatten progressive statements while preserving simple‑present ones, and this asymmetry persists across multiple configurations. Experiments reveal that the loss of temporal information can alter how downstream readers assess the validity of a fact and influence their decision‑making, sometimes leading them to act without further verification.

By Sugam Panthi, Muhaiminul Yeamin, Siyan Luo, Rabab Abdelfattah
arXiv Machine Learning
Sep 4

MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval

MemoryLACE (MemLACE) is a lightweight memory framework that explicitly models the lifecycle of textual evidence—capturing sparse merge, supersession, and contradiction relations—while preserving atomic natural‑language memories and their provenance. Unlike traditional systems that retrieve memories independently, MemLACE reconstructs relation‑aware evidence units that expose current, historical, supporting, and conflicting evidence for downstream reasoning. In benchmark evaluations (BEAM and StructMemEval) using both open‑weight and proprietary LLM backbones, MemLACE achieves the highest overall performance among same‑backbone comparisons and reduces BEAM runtime by 66.6% compared to the strongest reflective‑memory baseline, Hindsight.

By Meriem Yacoubi, Pia Schmidt, Nenad Petrovic, Ahmed Frikha, Martin Kirchhoff, Alois Knoll
arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi