How Much Memory Does Your Agent Actually Need?
Related stories
Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents
arXiv:2608.31057v1 Announce Type: new Abstract: Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles a...
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
arXiv:2607. 01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators.
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning.
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem
arXiv:2608. 11654v1 Announce Type: new Abstract: Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal.
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem
Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal. This paper takes a first step toward this account.
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
arXiv:2607. 13157v1 Announce Type: new Abstract: Agent memory is a systems problem for long-horizon agents.
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents
arXiv:2607. 08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and storing what happened in previous turns and sessions.
DolphinBench: Mapping the Pareto Frontier of Agent Memory
arXiv:2609.24971v1 Announce Type: new Abstract: Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are bu...
PM-Bench: Evaluating Prospective Memory in LLM Agents
arXiv:2607. 12385v1 Announce Type: new Abstract: A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing.
DolphinBench: Mapping the Pareto Frontier of Agent Memory
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format,...
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
arXiv:2605. 18421v2 Announce Type: replace-cross Abstract: Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution.