Planning with the Views
arXiv:2605. 29563v3 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2605. 29563v2 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2605. 29563v3 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when...
arXiv:2609.16737v1 Announce Type: cross Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches of...
CueNav is a video model-based navigation framework that uses visual cues—a Bird's-Eye View map for global task context and a body-aware egocentric view for embodiment context—to guide a video planner. The framework couples this planner with an embodiment-specific Inverse-Dynamics Model that translates dense flow fields from the video plan into robot actions. Experiments show that CueNav nearly doubles maze navigation success compared to cue-less planning and achieves 70% success in narrow passages, while also supporting zero-shot semantic-conditioned navigation across different robot platforms.
arXiv:2606. 17386v1 Announce Type: cross Abstract: End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments.
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guida...
arXiv:2607. 26005v1 Announce Type: cross Abstract: Self-play in simulation produces robust driving policies at scale.
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
arXiv:2608.29434v1 Announce Type: cross Abstract: JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction...
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
arXiv:2606. 14879v1 Announce Type: cross Abstract: Mobile agents require efficient exploration strategies to map unseen environments and autonomously plan tasks.