arXiv AI By Xijie Huang, Yongyang Wan, Chengbin Dong, Zimo Ding, Mo Zhu, Yijin Wang, Zhiyang Liu, Fei Gao, Yuze Wu, Xin Zhou

NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation

Read the original on arXiv AI →

NavGen introduces a text-to-video data generation pipeline that creates about 400K vision‑language navigation episodes for both indoor and outdoor scenes, using high‑fidelity visual generative models. The approach includes a style‑diversification method to scale up rare, hard‑to‑collect data. Models trained on NavGen data outperform those trained on existing UAV navigation datasets and achieve a 75% success rate in real‑world flying experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
3d ago

SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

SatNav is a new, scalable benchmark for long‑horizon vision‑language navigation (VLN) with unmanned aerial vehicles (UAVs), built from high‑resolution satellite imagery. It generates 118,000 navigation episodes across 59 scenes in 18 cities, using satellite crops to approximate UAV nadir views and featuring three task families—Boundary, Landmark, and Route—to test long‑term memory and geospatial reasoning. The benchmark also introduces SwiftVLN, a modular framework for memory component experimentation, and demonstrates that models trained on satellite data can transfer to real‑flight UAV observations.

By Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang, Jie Song, Xiao Hu