Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal. This paper takes a first step toward this account.
arXiv:2608. 11654v1 Announce Type: new Abstract: Despite the wide deployment of memory in large-model agents, there is no unified formal account of what a memory is or when it is optimal.
By Hongyao Tang
The paper discusses the agent-centric general value function (ACGVF) framework, which allows an agent to decide both which goal to pursue and when to consider a goal finished, beyond merely selecting actions. It notes that ACGVF assumes full observability, while a prior approach used an internal belief state but required externally supplied goals. The note proposes to unify and extend these methods using hierarchical hidden Markov models (HHMMs).
By Kevin Murphy
arXiv:2511. 22226v2 Announce Type: replace Abstract: The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled from their environment, such that policies are treated as being separate from the world they inhabit.
By Alexander Meulemans, Rajai Nasser, Maciej Wo{\l}czyk, Marissa A. Weis, Seijin Kobayashi, Blake Richards, Guillaume Lajoie, Angelika Steger, Marcus Hutter, James Manyika, Rif A. Saurous, Jo\~ao Sacramento, Blaise Ag\"uera y Arcas
The paper introduces Imagine-then-Plan (ITP), a framework that lets agents learn by interacting with a learned world model to generate multi-step imagined trajectories. ITP features an adaptive lookahead mechanism that balances ultimate goals with task progress, producing richer signals about future outcomes. Experiments on various benchmarks show that ITP outperforms existing baselines, and analyses suggest the adaptive lookahead improves reasoning for complex tasks.
By Youwei Liu, Jian Wang, Hanlin Wang, Beichen Guo, Wenjie Li
arXiv:2609.36595v1 Announce Type: cross
Abstract: Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual...
By Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang, Zijian Jin, Shilong Liu, Ding Zhao
The paper investigates how a bounded agent should allocate its limited memory and communication resources when making decisions. It defines the remembering–signaling frontier as the set of memory and message rate pairs that achieve a given performance threshold for a fixed task and decision rule. The authors hypothesize that when history can reduce task loss more, the agent will need less peer communication, and preliminary referential game experiments support this idea.
By Yashar Talebirad, Eden Redman, Ali Parsaee, Osmar R. Zaiane
The paper introduces Retrospective World Modeling, a new paradigm for vision‑language‑model (VLM) agents that allows them to reason backward by estimating which action most likely caused a state transition. It proposes the Self‑Consistency Reward (SCR), an intrinsic signal that measures how well a policy action aligns with this retrospective explanation, providing dense transition‑level feedback. Experiments demonstrate that incorporating SCR improves policy robustness and generalization compared to purely prospective world‑modeling approaches.
By Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
arXiv:2607. 26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causal mechanisms.
By Jasorsi Ghosh
The paper introduces epistemic memory, a validity-maintenance layer for intelligent systems that tracks when stored knowledge remains applicable. It formalizes a dynamic epistemic quotient and shows that fixed semantic representations inevitably incur error as epistemic boundaries shift. The authors propose Observable Belief Memory (OBM), which combines current epistemic quotients, belief over quotient classes, and within-class provenance, and demonstrate that explicit epistemic tracking improves robustness under changing observation conditions.
By Pin-Han Ho, Limei Peng, Yiming Miao, Yan Jiao
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agent's current decision.
arXiv:2608.30067v1 Announce Type: cross
Abstract: How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-m...
By Ruize Xu, Xiao Yu, Yujin Tang, Chenming Shang, Nikhil Singh