Language-Conditioned World Modeling for Visual Navigation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces UniWM, a unified, memory‑augmented world model that merges egocentric visual foresight and planning into a single multimodal autoregressive backbone. By grounding action selection in visually imagined outcomes and using a hierarchical memory to fuse short‑term perception with long‑term trajectory context, UniWM aligns prediction with control and improves navigation stability. Experiments on four challenging benchmarks and the 1X Humanoid Dataset show up to 30% higher success rates, reduced trajectory errors, zero‑shot generalization to unseen datasets, and scalability to high‑dimensional humanoid navigation.
arXiv:2609.37250v1 Announce Type: cross Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
arXiv:2607.10744v5 Announce Type: replace Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models...
arXiv:2609.39235v1 Announce Type: cross Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
arXiv:2606. 17924v1 Announce Type: cross Abstract: Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation.
The paper introduces Instruct-to-Act, a system that decouples high‑level planning from low‑latency control by combining a vision‑language model (VLM) planner with a world‑model controller. The VLM generates sparse, high‑level text instructions, while the controller executes them autonomously at high frequency. Experiments across seven embodied environments, including multi‑agent settings, show that this approach outperforms both controller‑only and direct VLM action‑generation methods, maintains fast control, and allows swapping in different pretrained VLM planners without fine‑tuning.