The paper investigates how long‑horizon language model agents encode memory‑management signals before taking actions. By examining hidden states just prior to each action, the authors find that the model already signals the need for compression and recall, independent of context length or interaction progress, and that these signals vary across model depth. They propose the Preaction Memory with Evidence Retrieval (PaMER) framework, which uses state‑guided compression and selective evidence retrieval to reduce context consumption while preserving task performance.
By Mingxuan Wang, Guorun Yao, Fei Luo, Yinglong Guo, Chao Ning, Bo Wang, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv:2510. 00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise records of actions and observations.
By Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A. Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, Saravan Rajmohan
Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. Unlike static Chain of Thought compression, removing historical reasoning can change future actions and the resulting interaction trajectory.
Long‑horizon language model agents accumulate reasoning history, which inflates context length and inference cost. The paper introduces Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training‑free online method that ranks and removes reasoning blocks based on frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR raises average reward from 0.699 to 0.718 and cuts input, output, and cache read tokens by 25.5%, 14.4%, and 33.3% respectively, while analyses show that historical reasoning becomes replaceable once task‑relevant state is externalized.
By Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv:2608. 02515v1 Announce Type: cross Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context.
By Zhichen Liu, Ruihan Sun, Hengjie Yang, Zipeng Wu, Zhaohan Chen, Xiaofan Zhang, Yang Xu
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States introduces PoS, an inference-time framework that builds and maintains explicit belief states to guide large language model agents. Each belief state combines an estimate of the current world with unresolved task requirements, making clear what the agent still needs to learn and accomplish. PoS validates consistency, monitors task progress to detect Belief Trapping, and tailors recovery to the trapping pattern and unresolved requirements, achieving top performance across four benchmarks with all three LLM backbones.
By Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen, Rongchen Gao, Yongqian Sun, Shenglin Zhang, Guiyang Liu, Cheng Zhang, Fang Situ, Qi Zhou, Dan Pei