Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory
arXiv:2606. 10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions.
The paper "Beyond Static Summarization: Proactive Memory Extraction for LLM Agents" identifies two shortcomings in current memory extraction for large language model agents: (1) extraction occurs ahead of time and mixes multiple types of information, leading to loss of useful details, and (2) extraction is typically one‑off, allowing errors and hallucinations to persist. To address these issues, the authors propose ProMem, a proactive framework that separates details, events, and relations, applies distinct extraction strategies for each, checks for completeness, and verifies facts at an atomic level. Experiments demonstrate that ProMem enhances memory completeness and question‑answering accuracy while maintaining a favorable balance between quality and token cost.
arXiv:2606. 10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions.
DyadMem introduces a new benchmark for evaluating how long‑term agents remember user‑specific relational information, called User‑conditioned Relational Agent Memory (URAM). The dataset contains 3,065 episodes, 50,961 sessions, and 61,210 QA pairs, with detailed annotations for memory capture, update, and recall across multi‑session trajectories. Experiments on 20 models show strong performance in a gold‑memory setting but a significant drop in full‑pipeline QA, highlighting gaps in current LLM memory capabilities.
arXiv:2610.02472v1 Announce Type: cross Abstract: Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (...
arXiv:2607. 16211v1 Announce Type: new Abstract: LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment.
arXiv:2608. 03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history.
The paper introduces the Weighted Memory Tree (WMT), a hierarchical memory system for large language model agents that organizes execution histories into tasks, subtasks, and actions while assigning each memory a dynamic retention score. Event-based updates and selection-based decay allow WMT to preserve useful information, fold completed trajectories, suppress low-utility content, and retain access to folded context. Experiments on GAIA-Text with Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B show that WMT improves accuracy by an average of 9.97 percentage points and reduces prompt-token usage by 32.8%, while also limiting the persistence of unreliable information.
UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.
Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems have made past experiences more persistent, compact, and retrievable, but retrieval alone does not ensure that a memory provides valid evidence for the current query.
arXiv:2606. 06787v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge.
arXiv:2507. 05257v4 Announce Type: replace-cross Abstract: Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, while another critical component-memory, encompassing how agents memorize, update, and retrieve long-term information-is under-evaluated due to the lack of benchmarks.
The paper introduces StructMemEval, a benchmark designed to assess how well large language model (LLM) agents can organize their long‑term memory rather than merely recall facts. It compiles tasks that humans typically solve by structuring knowledge—such as transaction ledgers, to‑do lists, and trees—and evaluates agents on these. Experiments show that simple retrieval‑augmented LLMs struggle with such organization tasks, while memory‑augmented agents perform better when explicitly prompted to structure their memory, yet many modern LLMs still fail to recognize memory structures without prompting.
EnSIMem is an entity‑structured long‑term memory architecture designed for agents that interact with users over extended periods. It organizes interactions into theme‑coherent episodes and creates dialogue‑grounded index entries of the form [entity][entity type][property:value], preserving source turns, temporal data, and multimodal fields. During online interaction, the agent decomposes requests into evidence requirements, performs entity‑property lookup, and retrieves the necessary evidence to generate responses directly from preserved source material rather than lossy summaries.