arXiv:2607. 25066v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window.
By Thang Dang, Yuma Ichikawa, Sakina Fatima, Koichi Shirahata
The paper introduces Just-in-Time Memory (JitMem), a system that defers memory curation until a task is read, allowing a curator to synthesize task‑specific memory payloads based on the current query. Unlike traditional write‑time curation, JitMem retains raw trajectories and trains the curator using immediate task success, avoiding long‑horizon credit‑assignment issues. Experiments on ALFWorld, WebShop, and τ²‑bench show JitMem consistently outperforms both no‑memory agents and existing write‑time memory methods, with improvements of up to 16.3 absolute success‑rate points.
whyItMatters":"By curating memory at read time, JitMem enables more effective, task‑adaptive recall that directly improves agent performance across diverse benchmarks."
By Yefan Zhou, Yang Li, Zeyu Leo Liu, Semih Yavuz, Shafiq Joty
arXiv:2607. 10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments.
By Yixiong Chen, Xinyi Bai, Alan Yuille
arXiv:2609.35808v1 Announce Type: cross
Abstract: Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current exe...
By Quanquan Li, Hongbo Zhang, Yihe Chi, Liuyang Song, Jingyu Li, Yuxiang Huang, Hongzhen Zhang, Guitao Cao
arXiv:2609. 04875v1 Announce Type: cross Abstract: Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache.
By Chao Yao, Yangbo Wei, Zhen Huang, Junhong Qian, Chenle Chen, Shaoqiang Lu, Chen Wu, Lei He
The paper introduces MERIT, a benchmark that evaluates the marginal benefit of long‑term memory for tool‑using large language model agents while explicitly accounting for cost. MERIT provides episodic tool‑use tasks across three domains, verifies dependence on earlier‑episode facts, and measures memory operations in tokens and dollars. Experiments on GPT‑4.1‑mini, Claude Haiku 4.5, and Claude Sonnet 5 show that memory can significantly improve task success, but its utility varies widely across models and memory implementations, and full replay is rarely cost‑effective.
By Shweta Mishra, Shashank Mishra