Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning cap...
arXiv:2604. 25416v2 Announce Type: replace Abstract: Model-based reinforcement learning distinguishes between dynamics models operating on proprioceptive states and latent dynamics models typically operating on high-dimensional image observations.
The paper introduces Dual-Frontier, a learning principle that determines when an agent should trust its world model for decision-making. It formalizes the failure-attribution problem as a counterfactual decomposition of return loss and shows that its components cannot be identified from passive interaction, even for finite-horizon planners. Dual-Frontier allows a model‑guided decision only when the predicted advantage exceeds a certified bound on decision‑relevant world‑model error; otherwise, the agent focuses on verifying the model. The authors provide theoretical guarantees, adaptive evidence reuse, and experimental validation on controlled and realistic benchmarks, demonstrating improved decision quality and reliability.
arXiv:2606. 31422v2 Announce Type: replace Abstract: Language agents acting over long horizons must maintain beliefs about tool states, object locations, graph edges, and subgoal dependencies.
The paper introduces Imagine-then-Plan (ITP), a framework that lets agents learn by interacting with a learned world model to generate multi-step imagined trajectories. ITP features an adaptive lookahead mechanism that balances ultimate goals with task progress, producing richer signals about future outcomes. Experiments on various benchmarks show that ITP outperforms existing baselines, and analyses suggest the adaptive lookahead improves reasoning for complex tasks.
arXiv:2609.05834v1 Announce Type: new Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the...