When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict
arXiv:2608. 13921v1 Announce Type: new Abstract: LLM agents increasingly maintain personal memory across sessions, but it can conflict.
arXiv:2608. 13921v1 Announce Type: new Abstract: LLM agents increasingly maintain personal memory across sessions, but it can conflict.
arXiv:2607. 10526v1 Announce Type: new Abstract: Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills.
arXiv:2601. 09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the model's parametric knowledge.
arXiv:2607. 12893v1 Announce Type: new Abstract: Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions.
arXiv:2605.12978v2 Announce Type: replace Abstract: Learning from past experience benefits from two complementary forms of memory: episodic traces -- raw trajectories of what happened -- and consolid...
arXiv:2608. 08236v1 Announce Type: new Abstract: Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be trusted.
The paper introduces FAME, a training‑free framework that evaluates false memory in autonomous agents by tracking how their internal beliefs shift under counterfactual scenarios. False memory, defined as biases arising from spurious correlations, environment shifts, or knowledge conflicts, is hard to detect with standard methods. FAME measures concept drift in hidden states, achieving AUROCs between 76.2% and 96.7% and outperforming baselines on benchmarks such as GSM‑Symbolic, GitChameleon, and BigBench‑Hard.
UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.
arXiv:2608. 07438v1 Announce Type: new Abstract: Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible.
arXiv:2609. 04875v1 Announce Type: cross Abstract: Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache.
arXiv:2607. 10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments.
arXiv:2606. 23195v2 Announce Type: replace Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence.