arXiv:2607. 01767v1 Announce Type: new Abstract: As agent planning moves from short tool chains toward persistent workflows with thousands or tens of thousands of steps, failures will occur inside large planning graphs rather than in isolated predictions.
By Xinyuan Song, Zekun Cai
arXiv:2606. 27806v1 Announce Type: new Abstract: World models for language agents come in two useful forms.
By Xinyuan Song, Zekun Cai
The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.
By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
The paper introduces the Graph Dynamics Model (GDM), a world model that learns stochastic latent dynamics over evolving graph topologies. GDM employs a sparse recurrent adjacency matrix for topology updates and a recurrent state‑space architecture for stochastic transitions, enabling it to handle partially observable, stochastic environments. The authors also propose the Graph Distribution Distance (GDD) metric, using maximum mean discrepancy with a graph kernel, to compare predicted and true joint graph state distributions, and demonstrate GDM’s superior performance and zero‑shot generalisation on large graphs.
By Alex Schutz, Nick Hawes, Victor-Alexandru Darvariu
arXiv:2606. 27806v3 Announce Type: replace Abstract: Language agents plan by generating not only actions but also implicit predictions of how the world will change.
By Xinyuan Song, Zekun Cai
arXiv:2607. 08894v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during planning, leading to high computational costs and stochastic behavior.
By Maureese Williams, Dymitr Nowicki
arXiv:2510. 09416v4 Announce Type: replace Abstract: Learning on temporal graphs has become a central topic in graph representation learning, with numerous benchmarks indicating the strong performance of state-of-the-art models.
By Abigail J. Hayes, Tobias Schumacher, Markus Strohmaier
arXiv:2608.24855v1 Announce Type: new
Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
By Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang
arXiv:2608. 08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error.
By Wei Wang, Yaosen Chen, Han Yang, Yuegen Liu, Mingli Luo, Xinxin Jiao, Xuming Wen, Ming Liu
arXiv:2607. 15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation.
By Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang
GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.
By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic
Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational overhead. Reusing a cached plan reduces this overhead, yet its effectiveness depends on how prediction mismatch propagates through the local dynamics.