arXiv AI

When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense

arXiv:2606. 22030v2 Announce Type: replace Abstract: We investigate when belief-based memory actually improves large language model (LLM) agents.

arXiv Machine Learning
Sep 3

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

The paper introduces CAPTURE, a system designed to help personalized language agents distinguish genuine preference changes from temporary context shifts or malicious memory poisoning. CAPTURE employs a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing to resolve ambiguity. Experiments on 480 episodes from 96 users show CAPTURE outperforms baseline methods, limiting poisoning success while accepting most real preference updates.

By S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol
arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi
arXiv Computation and Language
Sep 1

Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents

Agent Zero Memory is a provenance‑aware long‑term memory system for large language model agents that distills user interactions into three parallel memory structures: an episodic timeline, an associative entity‑event knowledge graph, and a semantic, citation‑locked hierarchical documentary memory. Retrieval is performed via an intent gate, source router, and concurrent searches across the three systems, producing integrated, cited answers that exclude fabrication and require evidence the reader has opened. The system achieves state‑of‑the‑art performance on LongMemEval (95.60%) and LoCoMo (93.60%) while offering a favorable accuracy‑cost‑latency trade‑off across multiple backbone LLMs.

By Ming Wu, Pengyuan Zhu
arXiv AI
Sep 10

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.

By Shweta Mishra, Shashank Mishra
arXiv AI
3d ago

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

The paper introduces FAME, a training‑free framework that evaluates false memory in autonomous agents by tracking how their internal beliefs shift under counterfactual scenarios. False memory, defined as biases arising from spurious correlations, environment shifts, or knowledge conflicts, is hard to detect with standard methods. FAME measures concept drift in hidden states, achieving AUROCs between 76.2% and 96.7% and outperforming baselines on benchmarks such as GSM‑Symbolic, GitChameleon, and BigBench‑Hard.

By Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu