VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots
Read the original on arXiv AI →The paper introduces VLN on the Fly, an onboard vision‑language navigation stack for aerial robots that separates grounding, planning, and control into inspectable stages. A quantized vision‑language model grounds instructions to a coarse image cell, depth estimation lifts this to a 3D goal, a fast B‑spline planner generates a feasible trajectory, and a pretrained reinforcement learning policy translates the trajectory into motor commands. In controlled indoor flights, the stack achieved the target in 13 of 15 trials with a mean goal error of 5.72 cm and 39.3% GPU utilization, and successfully tracked collision‑free trajectories in cluttered environments.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.