arXiv Machine Learning

Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality

arXiv Machine Learning
Jul 14

A Control Theory of Predictability in Latent World Models

arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.

By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv AI
Sep 25

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM is a new action‑discriminative joint‑embedding world model designed for counterfactual model predictive control. It augments residual latent dynamics with action‑recovery regularization based on inverse dynamics and conditional mutual information, while discarding auxiliary heads at test time so that MPC remains unchanged. Experiments on OGBench‑Cube and other simulation environments show substantial gains in hard‑start success and mean success, and zero‑shot transfer to a Franka robot improves pick‑and‑place success from 42.2% to 71.1%.

By Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao
arXiv AI
3d ago

ATM: Why Latent World Models Can Fail to Plan

The paper investigates why latent world models, despite accurate latent predictions, can perform poorly in downstream planning. It introduces the concept of action-identifiability and formalizes it via Bayes inverse risk, showing that self-decodable predicted transitions may encode domain‑specific action relationships that do not transfer to real environments. Using the Action‑Consistency Transfer Matrix (ATM), the authors demonstrate that true‑transition action‑identifiability correlates strongly with planning success across several benchmarks, and that the ATM can diagnose cross‑domain inconsistencies and aid lightweight model screening.

By Jiaheng Chen, Tinghe Zhang, Yucheng Xiao, Xinyong Cai, Lan Yu, Juncheng Bu, Jiaxing Li, Yunlong Wang
arXiv Machine Learning
Jun 8

Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning

arXiv:2511. 02748v2 Announce Type: replace-cross Abstract: We argue that sixth-generation (6G) intelligence is not fluent token prediction but the capacity to imagine and choose -- to simulate future scenarios, weigh trade-offs, and act with calibrated uncertainty.

By Farhad Rezazadeh, Amir Ashtari Gargari, Hatim Chergui, Sandra Lagen, Merouane Debbah, Houbing Song, Lingjia Liu
arXiv AI
4d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv AI
Sep 25

Dual-Frontier: When Can an Agent Trust Its World Model?

The paper introduces Dual-Frontier, a learning principle that determines when an agent should trust its world model for decision-making. It formalizes the failure-attribution problem as a counterfactual decomposition of return loss and shows that its components cannot be identified from passive interaction, even for finite-horizon planners. Dual-Frontier allows a model‑guided decision only when the predicted advantage exceeds a certified bound on decision‑relevant world‑model error; otherwise, the agent focuses on verifying the model. The authors provide theoretical guarantees, adaptive evidence reuse, and experimental validation on controlled and realistic benchmarks, demonstrating improved decision quality and reliability.

By Huatai Zhu, Qiang Chen, Ziqian Kou, Wenhao Li, Fei Wang, Yichao Cao, Xiu Su, Yi Chen
arXiv Computation and Language
Sep 14

LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?

LLM‑BabyBench transforms the BabyAI gridworld into a fully observable, purely textual setting that isolates planning as the sole source of failure. By serialising the entire grid, providing formal instructions, and validating actions deterministically, the benchmark introduces the PPD suite—Predict, Plan, and Decompose tasks—each scored with metrics that separate mission understanding from sequencing. Across a range of large language models, simulation accuracy is high while planning success drops sharply beyond a model‑specific horizon, revealing that plan length—not grid size—drives failure and that models often commit to a single corridor‑shaped route without backtracking.

By Idriss Malek, Omar Choukrani, Daniil Orel, Anh Duy Le Dinh, Zhuohan Xie, Zangir Iklassov, Martin Tak\'a\v{c}, Salem Lahlou