arXiv AI

Understanding Rollout Error in Graph World Models

arXiv:2606. 27780v1 Announce Type: new Abstract: World models are often used for planning by rolling learned dynamics forward.

arXiv Machine Learning
1d ago

Learning Commute-Time-Preserving World Models for Planning

The paper introduces Commute-Time-Preserving World Models (CTWMs), which learn latent representations that reflect commute-times in an environment by using a latent displacement predictor and a log-determinant regularizer. This approach addresses the issue that existing self-supervised methods degrade the necessary eigenvalue-dependent scaling for accurate commute-time representation. In experiments, CTWMs outperform the task-agnostic baseline LeWM on several continuous goal-reaching benchmarks while using only half the parameters.

By Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch, Friedemann Zenke
arXiv AI
Sep 25

Beyond Static Graph World Models: Learning Stochastic Latent Dynamics over Evolving Topologies

The paper introduces the Graph Dynamics Model (GDM), a world model that learns stochastic latent dynamics over evolving graph topologies. GDM employs a sparse recurrent adjacency matrix for topology updates and a recurrent state‑space architecture for stochastic transitions, enabling it to handle partially observable, stochastic environments. The authors also propose the Graph Distribution Distance (GDD) metric, using maximum mean discrepancy with a graph kernel, to compare predicted and true joint graph state distributions, and demonstrate GDM’s superior performance and zero‑shot generalisation on large graphs.

By Alex Schutz, Nick Hawes, Victor-Alexandru Darvariu
arXiv Machine Learning
Jul 17

What Do Temporal Graph Learning Models Learn?

arXiv:2510. 09416v4 Announce Type: replace Abstract: Learning on temporal graphs has become a central topic in graph representation learning, with numerous benchmarks indicating the strong performance of state-of-the-art models.

By Abigail J. Hayes, Tobias Schumacher, Markus Strohmaier
arXiv AI
Jul 20

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

arXiv:2607. 15610v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sampled complete trajectories for advantage estimation.

By Xintong Li, Sha Li, Yuwei Zhang, Changlong Yu, Rongmei Lin, Hongye Jin, Shuyi Guan, Xin Liu, Linwei Li, Qingyu Yin, Jingbo Shang
arXiv AI
Sep 18

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL is a framework that uses an explicit graph world model to verify and repair long‑horizon plans generated by large language models (LLMs). The graph encodes object relations, action pre‑conditions and effects, and probabilistic beliefs about unobserved object locations, allowing the system to predict action outcomes, detect violations, and repair them before execution. In experiments on BEHAVIOR‑1K, GAVEL boosts single‑task success from 41.2 % to 91.8 % and multi‑task success from 19.9 % to 92.6 %, while also reducing travel distance by about 5.4 % compared with a static variant.

By Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic