Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
CueNav is a video model-based navigation framework that uses visual cues—a Bird's-Eye View map for global task context and a body-aware egocentric view for embodiment context—to guide a video planner. The framework couples this planner with an embodiment-specific Inverse-Dynamics Model that translates dense flow fields from the video plan into robot actions. Experiments show that CueNav nearly doubles maze navigation success compared to cue-less planning and achieves 70% success in narrow passages, while also supporting zero-shot semantic-conditioned navigation across different robot platforms.
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guida...
arXiv:2606. 29879v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning.
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations.
arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when...