APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
arXiv:2606. 03441v1 Announce Type: cross Abstract: Autonomous vision-based perching of quadrotors on moving inclined platforms is critical for air-ground collaboration but remains challenging due to the limited field of view (FOV).
The paper introduces VLN on the Fly, an onboard vision‑language navigation stack for aerial robots that separates grounding, planning, and control into inspectable stages. A quantized vision‑language model grounds instructions to a coarse image cell, depth estimation lifts this to a 3D goal, a fast B‑spline planner generates a feasible trajectory, and a pretrained reinforcement learning policy translates the trajectory into motor commands. In controlled indoor flights, the stack achieved the target in 13 of 15 trials with a mean goal error of 5.72 cm and 39.3% GPU utilization, and successfully tracked collision‑free trajectories in cluttered environments.
arXiv:2606. 03252v1 Announce Type: cross Abstract: Navigating a drone in unseen and cluttered environments requires reliable generalization to unseen scene layouts and understanding of environmental structure relative to the robot's capabilities.
arXiv:2606. 18634v1 Announce Type: cross Abstract: To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots.
arXiv:2601. 15995v2 Announce Type: replace-cross Abstract: Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion.
arXiv:2606. 14772v1 Announce Type: cross Abstract: Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions.
arXiv:2609.16737v1 Announce Type: cross Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches of...
arXiv:2510. 06277v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (GCRL) offers a unified way to pursue diverse tasks, yet most existing methods rely on state- or position-based goal representations that are unavailable in real-world robotics.
arXiv:2606. 15896v1 Announce Type: cross Abstract: Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective.
arXiv:2606. 03963v1 Announce Type: cross Abstract: Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks.
CueNav is a video model-based navigation framework that uses visual cues—a Bird's-Eye View map for global task context and a body-aware egocentric view for embodiment context—to guide a video planner. The framework couples this planner with an embodiment-specific Inverse-Dynamics Model that translates dense flow fields from the video plan into robot actions. Experiments show that CueNav nearly doubles maze navigation success compared to cue-less planning and achieves 70% success in narrow passages, while also supporting zero-shot semantic-conditioned navigation across different robot platforms.