Grounded World Model: Latent Planning with Language Goals
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
FIRM-WM is a compact pixel world model that separates a goal‑comparable configuration from a 128‑dimensional dynamic fiber, enabling reward‑free visual planning from offline videos. It addresses two key mismatches: aligning planning states with goal images and reconciling factual trajectories with interventional sampling. In experiments, FIRM‑WM achieves high success rates on TwoRoom, Reacher, and OGBench‑Cube while using fewer parameters and faster planning times than prior models.
arXiv:2606. 17924v1 Announce Type: cross Abstract: Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation.
arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
JEPA-WAM enhances World Action Models (WAMs) by pairing text instructions with stochastically generated visual cues, using a text-to-image generator and a frozen V‑JEPA encoder to create dense goal representations. These representations are compressed into goal tokens that condition both video and action experts via cross‑attention, enabling the model to better ground instructions. On a new real‑robot benchmark, JEPA‑WAM attains 87.3%, 74.5%, and 80.9% success rates across in‑distribution, out‑of‑distribution scenes, and out‑of‑distribution instructions, outperforming prior methods by significant margins.