arXiv AI By Zhishang Xiang, Zerui Chen, Yunbo Tang, Zhimin Wei, Ruqin Ning, Yujie Lin, Qinggang Zhang, Jinsong Su

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Read the original on arXiv AI →

arXiv:2607. 01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 1

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning.

arXiv AI
6d ago

Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

The paper introduces MemProbe, a framework inspired by cognitive science to evaluate stability-plasticity tradeoffs in agent memory systems. It offers four experimental paradigms—interference, misinformation, consolidation strength, and reconsolidation window—to manipulate memory updates, preservation, and uncertainty. Using a 56-episode diagnostic suite, the authors evaluate six incremental memory systems, revealing that similar overall scores can mask distinct behavioral profiles in how memories are updated, preserved, attributed, and temporally organized.

By Jiaqi Ding, Guorong Wu
arXiv Machine Learning
Sep 11

Evaluating Memory Structure in LLM Agents

The paper introduces StructMemEval, a benchmark designed to assess how well large language model (LLM) agents can organize their long‑term memory rather than merely recall facts. It compiles tasks that humans typically solve by structuring knowledge—such as transaction ledgers, to‑do lists, and trees—and evaluates agents on these. Experiments show that simple retrieval‑augmented LLMs struggle with such organization tasks, while memory‑augmented agents perform better when explicitly prompted to structure their memory, yet many modern LLMs still fail to recognize memory structures without prompting.

By Alina Shutova, Alexandra Olenina, Ivan Vinogradov, Anton Sinitsin