Controlled Memory Interference in Continual LLM Agents
arXiv:2608. 07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience.
The paper introduces MemProbe, a framework inspired by cognitive science to evaluate stability-plasticity tradeoffs in agent memory systems. It offers four experimental paradigms—interference, misinformation, consolidation strength, and reconsolidation window—to manipulate memory updates, preservation, and uncertainty. Using a 56-episode diagnostic suite, the authors evaluate six incremental memory systems, revealing that similar overall scores can mask distinct behavioral profiles in how memories are updated, preserved, attributed, and temporally organized.
arXiv:2608. 07622v1 Announce Type: new Abstract: Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience.
arXiv:2607. 01071v1 Announce Type: cross Abstract: Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators.
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved memories often induce a critical issue of sycophancy, causing agents to over-align with the user at the cost of factual accuracy or objective reasoning.
arXiv:2607. 10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments.
arXiv:2606.24595v2 Announce Type: replace Abstract: Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that inter...
arXiv:2607. 17621v1 Announce Type: new Abstract: Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections.
arXiv:2606. 02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
arXiv:2606. 02461v2 Announce Type: replace Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future episodes.
arXiv:2609.21533v1 Announce Type: new Abstract: LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusin...
The paper re‑evaluates memory‑based self‑improving agents by adding multiple runs to measure variance and by randomizing task order. It finds that agent performance is noisy in complex, multi‑step environments and that improvement depends heavily on the sequence of tasks, revealing a hidden curriculum effect. The authors suggest that underspecification of tasks and environments contributes to this fragility and demonstrate that adding detailed rubrics and feedback can partially mitigate performance drops, though gaps remain.
arXiv:2602. 06052v4 Announce Type: replace-cross Abstract: Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-world evaluation.
The paper re‑evaluates memory‑based self‑improving agents by running multiple trials and shuffling task orders, revealing that agent performance is noisy and highly sensitive to task sequencing. It shows that implicit curricula in default task orders act as hidden prerequisites for success, and that underspecification of tasks and environments contributes to fragility. Adding detailed rubrics and environment feedback partially mitigates performance drops but significant gaps remain, underscoring the need for stricter evaluation protocols and better human oversight.