arXiv AI By Nikolai Smolyanskiy

Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander

Read the original on arXiv AI →

arXiv:2607. 01736v1 Announce Type: cross Abstract: We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 2

Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander

We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choosing the right checkpoint from a world-model training run is difficult: validation loss and multi-step prediction RMSE keep improving long after closed-loop performance has collapsed.

arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
Sep 25

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM is a new action‑discriminative joint‑embedding world model designed for counterfactual model predictive control. It augments residual latent dynamics with action‑recovery regularization based on inverse dynamics and conditional mutual information, while discarding auxiliary heads at test time so that MPC remains unchanged. Experiments on OGBench‑Cube and other simulation environments show substantial gains in hard‑start success and mean success, and zero‑shot transfer to a Franka robot improves pick‑and‑place success from 42.2% to 71.1%.

By Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao