arXiv AI

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

arXiv:2607. 01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators.

Hugging Face Trending Papers
Jul 1

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning.

arXiv AI
6d ago

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

The paper introduces MemProbe, a framework inspired by cognitive science to evaluate stability-plasticity tradeoffs in agent memory systems. It offers four experimental paradigms—interference, misinformation, consolidation strength, and reconsolidation window—to manipulate memory updates, preservation, and uncertainty. Using a 56-episode diagnostic suite, the authors evaluate six incremental memory systems, revealing that similar overall scores can mask distinct behavioral profiles in how memories are updated, preserved, attributed, and temporally organized.

By Jiaqi Ding, Guorong Wu
arXiv Machine Learning
Sep 11

Evaluating Memory Structure in LLM Agents

The paper introduces StructMemEval, a benchmark designed to assess how well large language model (LLM) agents can organize their long‑term memory rather than merely recall facts. It compiles tasks that humans typically solve by structuring knowledge—such as transaction ledgers, to‑do lists, and trees—and evaluates agents on these. Experiments show that simple retrieval‑augmented LLMs struggle with such organization tasks, while memory‑augmented agents perform better when explicitly prompted to structure their memory, yet many modern LLMs still fail to recognize memory structures without prompting.

By Alina Shutova, Alexandra Olenina, Ivan Vinogradov, Anton Sinitsin
arXiv AI
Jun 30

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

arXiv:2507. 05257v4 Announce Type: replace-cross Abstract: Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, while another critical component-memory, encompassing how agents memorize, update, and retrieve long-term information-is under-evaluated due to the lack of benchmarks.

By Yuanzhe Hu, Yu Wang, Julian McAuley
arXiv Computation and Language
Sep 21

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.

By Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland
arXiv Computation and Language
Aug 31

What Makes Agent Memory Useful for Reliable Unanswerable Question Handling?

The paper investigates how agent memory contributes to reliable handling of unanswerable questions (UAQs) within a unified Retrieval-Augmented Generation (RAG) framework. Four memory methods were evaluated across three UAQ datasets and two base models, revealing that memory can improve UAQ performance in selective settings but the gains are fragile under dataset shift. Procedural and rule-based memories, especially when combined with complementary behavioral signals, provide the most reliable support, indicating that effective UAQ memory relies more on transferable behavioral guidance than on sheer volume of stored experience.

By Chuanyuan Tan, Junjie Yu, Yuxin Wang, Yining Zheng, Xipeng Qiu, Wenliang Chen