Planning with the Views via Scene Self-Exploration
arXiv:2605. 29563v2 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2605. 29563v3 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2605. 29563v2 Announce Type: replace Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead?
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when...
arXiv:2609.16737v1 Announce Type: cross Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches of...
CueNav is a video model-based navigation framework that uses visual cues—a Bird's-Eye View map for global task context and a body-aware egocentric view for embodiment context—to guide a video planner. The framework couples this planner with an embodiment-specific Inverse-Dynamics Model that translates dense flow fields from the video plan into robot actions. Experiments show that CueNav nearly doubles maze navigation success compared to cue-less planning and achieves 70% success in narrow passages, while also supporting zero-shot semantic-conditioned navigation across different robot platforms.
Think3D introduces a framework that endows Vision‑Language Models with interactive 3D chain‑of‑thought reasoning by integrating 3D manipulation tools for active spatial exploration. The approach improves performance on benchmarks such as BLINK Multi‑view, MindCube‑1K, and VSI‑Bench‑Tiny for proprietary models like GPT‑4.1 and Gemini 2.5 Pro, and a reinforcement‑learning variant, Think3D‑RL, enables open‑weight models such as Qwen3‑VL‑4B to autonomously learn effective 3D exploration strategies, yielding tool‑use patterns comparable to stronger models and turning a performance drop on MindCube‑1K into a substantial improvement.
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guida...
arXiv:2607. 26005v1 Announce Type: cross Abstract: Self-play in simulation produces robust driving policies at scale.
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
arXiv:2606. 17386v1 Announce Type: cross Abstract: End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments.