APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2603. 25937v2 Announce Type: replace-cross Abstract: Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations.
NavGen introduces a text-to-video data generation pipeline that creates about 400K vision‑language navigation episodes for both indoor and outdoor scenes, using high‑fidelity visual generative models. The approach includes a style‑diversification method to scale up rare, hard‑to‑collect data. Models trained on NavGen data outperform those trained on existing UAV navigation datasets and achieve a 75% success rate in real‑world flying experiments.
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.
UrbanGround is a sandbox that tests how well multimodal large language model agents can translate local street‑view perception into reliable action within a physically realistic replica of Hong Kong. The platform offers closed‑loop first‑person interaction and an interactive map, allowing agents to navigate the 3D city and answer spatial questions. The study evaluates agents across three research questions—scene grounding, navigation over increasing distances, and robustness to route changes—revealing that while agents excel at visual recognition and short‑range reasoning, they struggle with sustained goal‑directed behavior and pedestrian‑aware movement.
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
arXiv:2607. 10991v1 Announce Type: cross Abstract: As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust.
arXiv:2607. 09792v1 Announce Type: cross Abstract: Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prior assumptions, limiting their robustness in open and uncertain real-world environments.
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
arXiv:2606. 07244v1 Announce Type: cross Abstract: Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments.
arXiv:2512. 19178v2 Announce Type: replace-cross Abstract: Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robotics.