arXiv Machine Learning By Zhichen Liu, Ruihan Sun, Hengjie Yang, Zipeng Wu, Zhaohan Chen, Xiaofan Zhang, Yang Xu

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Read the original on arXiv Machine Learning →

arXiv:2608. 02515v1 Announce Type: cross Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Memory Control Signals Emerge Before Action in Long Horizon Agents

The paper investigates how long‑horizon language model agents encode memory‑management signals before taking actions. By examining hidden states just prior to each action, the authors find that the model already signals the need for compression and recall, independent of context length or interaction progress, and that these signals vary across model depth. They propose the Preaction Memory with Evidence Retrieval (PaMER) framework, which uses state‑guided compression and selective evidence retrieval to reduce context consumption while preserving task performance.

By Mingxuan Wang, Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi