arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
By Alex Kwon
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it emits that stale value as a confident answer; give the same model an empty memory and it abstains.
arXiv:2609.37125v1 Announce Type: new
Abstract: Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition c...
By Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu, Bo Dong
arXiv:2609. 04875v1 Announce Type: cross Abstract: Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache.
By Chao Yao, Yangbo Wei, Zhen Huang, Junhong Qian, Chenle Chen, Shaoqiang Lu, Chen Wu, Lei He
arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
By Jie Wu, Ming Gong, Feixiang Cheng, Qinqin Zhao
The paper introduces PlanFence, a dependency-scoped action‑validation protocol for distributed large language model (LLM) agent teams. PlanFence requires plans to cite the exact public records they rely on, and executors validate only those records that could affect the pending action, replanning or blocking if validation is incomplete. In 30 controlled live workflows, a freshness‑only executor always acted on obsolete plans, whereas PlanFence completed all tasks without invalid actions, demonstrating controlled safety and system‑cost benefits.
By Evan Chen, Shiqiang Wang, Christopher G. Brinton