arXiv:2606. 10677v1 Announce Type: new Abstract: Long-term LLM agents need persistent memory that can track changing facts and provide relevant evidence across sessions.
By Suozhao Ji, Baodong Wu, Zehao Wang, Lei Xia, Qingping Li, Ruisong Wang, Wenbo Ding, Zhenhua Zhu, Boxun Li, Guohao Dai, Yu Wang
DyadMem introduces a new benchmark for evaluating how long‑term agents remember user‑specific relational information, called User‑conditioned Relational Agent Memory (URAM). The dataset contains 3,065 episodes, 50,961 sessions, and 61,210 QA pairs, with detailed annotations for memory capture, update, and recall across multi‑session trajectories. Experiments on 20 models show strong performance in a gold‑memory setting but a significant drop in full‑pipeline QA, highlighting gaps in current LLM memory capabilities.
By Yifei Tao, Xinyu Zhong, Henry Hengyuan Zhao, Fanyi Wang, Tengda Guo, Wentao Qiu, Ying Wang, Liujian Tang
arXiv:2610.02472v1 Announce Type: cross
Abstract: Personalized LLM assistants must recover sparse evidence from long conversation histories across queries of varying complexity. We introduce APDMem (...
By Chin-Lun Fu, Anagha Kulkarni, Hong Ni, Behrouz Madahian
arXiv:2607. 16211v1 Announce Type: new Abstract: LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructured storage loses relational context needed for multi-hop and temporal reasoning, and reliance on expensive LLM-based classification makes them impractical for latency-sensitive deployment.
By Zicheng Zhao, Xinyang Guo, Luyao Lv, Menghan Wang, Ming Li, Shuaicheng Li
arXiv:2608. 03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history.
By Yuxin Liao, Le Wu, Min Hou, Hao Liu, Han Wu, Zishu Wang
The paper introduces the Weighted Memory Tree (WMT), a hierarchical memory system for large language model agents that organizes execution histories into tasks, subtasks, and actions while assigning each memory a dynamic retention score. Event-based updates and selection-based decay allow WMT to preserve useful information, fold completed trajectories, suppress low-utility content, and retain access to folded context. Experiments on GAIA-Text with Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B show that WMT improves accuracy by an average of 9.97 percentage points and reduces prompt-token usage by 32.8%, while also limiting the persistence of unreliable information.
By Quang Dao, Purvi Kathalkar, Kenneth Eaton