arXiv Machine Learning

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

arXiv:2608. 02515v1 Announce Type: cross Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context.

arXiv AI
Sep 24

Memory Control Signals Emerge Before Action in Long Horizon Agents

The paper investigates how long‑horizon language model agents encode memory‑management signals before taking actions. By examining hidden states just prior to each action, the authors find that the model already signals the need for compression and recall, independent of context length or interaction progress, and that these signals vary across model depth. They propose the Preaction Memory with Evidence Retrieval (PaMER) framework, which uses state‑guided compression and selective evidence retrieval to reduce context consumption while preserving task performance.

By Mingxuan Wang, Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi
arXiv AI
Jul 8

From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space

arXiv:2607. 05794v1 Announce Type: new Abstract: Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retrieval interfaces, making the model a consumer of pre-selected evidence.

By Yue Xu, Yutao Sun, Yihao Liu, Mengyu Zhou, Jiayi Qiao, Lu Ma, Kai Tang, Wenjie Wang, Xiaoxi Jiang, Guanjun Jiang
arXiv AI
Aug 24

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

The paper introduces the Weighted Memory Tree (WMT), a hierarchical memory system for large language model agents that organizes execution histories into tasks, subtasks, and actions while assigning each memory a dynamic retention score. Event-based updates and selection-based decay allow WMT to preserve useful information, fold completed trajectories, suppress low-utility content, and retain access to folded context. Experiments on GAIA-Text with Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B show that WMT improves accuracy by an average of 9.97 percentage points and reduces prompt-token usage by 32.8%, while also limiting the persistence of unreliable information.

By Quang Dao, Purvi Kathalkar, Kenneth Eaton
arXiv AI
4d ago

CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory

CoEM introduces a Commit-on-Evidence Memory system that learns when to compress source evidence into compact memory facts while preserving potentially useful excerpts verbatim in a pending set. The system uses a learned policy to decide whether to promote, retain, or discard each pending excerpt as new context arrives, and a frozen verifier ensures only supported facts are committed. Reinforcement learning trains this policy with step-level evidence rewards and final answer rewards, leading to consistent improvements in long-context reasoning, achieving 10.4–11.4 F1 points over the strongest baseline on 6,400-document inputs.

By Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li, Can Rong, Heye Huang