arXiv Machine Learning

Learning Implicit Causal World Models from Multi-Agent Demonstrations

arXiv:2607. 26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causal mechanisms.

arXiv AI
3d ago

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

The paper introduces Retrospective World Modeling, a new paradigm for vision‑language‑model (VLM) agents that allows them to reason backward by estimating which action most likely caused a state transition. It proposes the Self‑Consistency Reward (SCR), an intrinsic signal that measures how well a policy action aligns with this retrospective explanation, providing dense transition‑level feedback. Experiments demonstrate that incorporating SCR improves policy robustness and generalization compared to purely prospective world‑modeling approaches.

By Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
Hugging Face Trending Papers
Aug 13

A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics.

arXiv AI
Sep 10

Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

The paper introduces an action‑conditioned world‑modeling framework that turns Earth‑system simulator trajectories into training data for controllable state‑transition learning. By pretraining on naturally observed state changes as implicit action supervision and using masked response learning, the model can infer unobserved variables and learn coupled system dependencies. Experiments on ecosystem dynamics across six global regions demonstrate that the model maintains long‑horizon emulation accuracy while enabling structural interventions and coherent responses in coupled ecosystem‑cycle variables.

By Zhihao Wang, Ruichen Wang, Ruohan Li, Lei Ma, George Hurtt, Xiaowei Jia, Gengchen Mai, Shaowen Wang, Yiqun Xie
arXiv AI
Sep 11

Reinforcement Learning with Temporal-Logic-Based Causal Diagrams

The paper introduces Temporal-Logic-based Causal Diagrams (TL-CDs) for reinforcement learning tasks that involve temporally extended goals. TL-CDs encode causal relationships among environmental properties, complementing deterministic finite automata that model rewards. By leveraging TL-CDs, the authors design an RL algorithm that can predict expected rewards early, leading to significantly reduced exploration and faster convergence to optimal policies.

By Yash Paliwal, Rajarshi Roy, Jean-Rapha\"el Gaglione, Nasim Baharisangari, Daniel Neider, Xiaoming Duan, Ufuk Topcu, Zhe Xu
arXiv AI
Aug 28

Predicting Consequences and Reinforcing Navigation Policies with Latent World Models

The paper introduces a Latent World Model (LWM) for robot navigation that predicts action‑conditioned latent feature compatibility instead of reconstructing future observations. By exploiting the correlation between spatial proximity and latent feature similarity, the model evaluates action consequences directly in latent space and supports counterfactual training using sampled action sequences. The learned world model can supervise policy learning from unlabeled video and further improve policies via reinforcement learning entirely within the model, eliminating the need for action annotations and additional environment interaction.

By Zengmao Wang, Wei Gao, Shuhan Shen