Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States introduces PoS, an inference-time framework that builds and maintains explicit belief states to guide large language model agents. Each belief state combines an estimate of the current world with unresolved task requirements, making clear what the agent still needs to learn and accomplish. PoS validates consistency, monitors task progress to detect Belief Trapping, and tailors recovery to the trapping pattern and unresolved requirements, achieving top performance across four benchmarks with all three LLM backbones.
By Yu Luo, Jiamin Jiang, Yimin Zuo, Xidao Wen, Rongchen Gao, Yongqian Sun, Shenglin Zhang, Guiyang Liu, Cheng Zhang, Fang Situ, Qi Zhou, Dan Pei
arXiv:2607. 20064v1 Announce Type: new Abstract: Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents.
By Alexis Fox, Junlin Wang, Paul Rosu, Bhuwan Dhingra
CoEM introduces a Commit-on-Evidence Memory system that learns when to compress source evidence into compact memory facts while preserving potentially useful excerpts verbatim in a pending set. The system uses a learned policy to decide whether to promote, retain, or discard each pending excerpt as new context arrives, and a frozen verifier ensures only supported facts are committed. Reinforcement learning trains this policy with step-level evidence rewards and final answer rewards, leading to consistent improvements in long-context reasoning, achieving 10.4–11.4 F1 points over the strongest baseline on 6,400-document inputs.
By Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma, Xianjie Dai, Han Zheng, Benwang Chen, Li Li, Can Rong, Heye Huang
arXiv:2606. 03329v1 Announce Type: new Abstract: Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts.
By Tiancheng Han, Yong Li, Wuzhou Yu, Qiaosheng Zhang, Wenqi Shao
AgenticRag‑R1 is a reinforcement‑learning framework that integrates reasoning, retrieval, and memory through a stack and fine‑grained action space. It uses hierarchical action‑aware rewards and an information‑aware trajectory rejection strategy to support long‑horizon learning. Experiments on multi‑hop, open‑domain, and agentic reasoning benchmarks show that AgenticRag‑R1 outperforms strong baselines and produces robust, interpretable, memory‑aware reasoning behaviors.
By Xinke Jiang, Yue Fang, Zhibang Yang, Jiaran Gao, Zhixin Zhang, Tao Feng, Rihong Qiu, Wentao Zhang, Hongxin Ding, Ruizhe Zhang, Yongxin Xu, Yuheng Huang, Xu Chu, Junfeng Zhao, Yasha Wang
arXiv:2606. 13316v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs).
By Xucong Wang, Ziyu Ma, Yong Wang, Shidong Yang, Hailang Huang, Renda Li, Pengkun Wang, Xiangxiang Chu
arXiv:2510. 00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise records of actions and observations.
By Minki Kang, Wei-Ning Chen, Dongge Han, Huseyin A. Inan, Lukas Wutschitz, Yanzhi Chen, Robert Sim, Saravan Rajmohan
arXiv:2606. 22030v2 Announce Type: replace Abstract: We investigate when belief-based memory actually improves large language model (LLM) agents.
By Pranav Singh
StateComp introduces a method for long‑horizon agents to decide when to compress historical interactions based on the current agent state, rather than relying on fixed windows or periodic schedules. The framework uses a two‑stage annotation process to create KEEP and READY labels, trains an imbalance‑aware router on frozen language model representations, and groups adjacent READY interactions into compact summaries. Experiments on WorkBuddyBench show that StateComp cuts agent and summarization tokens by 52.27% and speeds up representation extraction 12.67‑fold while preserving task performance.
By Mingxuan Wang, Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang, Guorun Yao, Yanbiao Ma, Jungong Han
arXiv:2609.37930v1 Announce Type: cross
Abstract: Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a cred...
By Jiaming Tang, Mingyan Liu, Armin Sarabi
arXiv:2606. 31650v1 Announce Type: cross Abstract: Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows.
By Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen
arXiv:2605.30219v2 Announce Type: replace
Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and...
By Haoming Xu, Weihong Xu, Zongrui Li, Mengru Wang, Yunzhi Yao, Chiyu Wu, Jin Shang, Yu Gong, Shumin Deng