arXiv:2609.08442v1 Announce Type: cross
Abstract: Aerial Vision-and-Language Navigation requires drones to follow natural-language instructions and navigate through complex urban environments. Accura...
By Shanwei Fan, Bin Zhang, Zhiwei Xu, Yingxuan Teng, Siqi Dai, Lin Cheng, Guoliang Fan
arXiv:2606. 30576v1 Announce Type: cross Abstract: Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.
By Liyao Wang, Ruipu Wu, Haojun Xu, Lei Shi, Linjiang Huang, Si Liu
arXiv:2606. 06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation.
By Shengtao Zheng, Kai Li, Weichen Zhang, Yu Meng, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
GS‑VLA introduces a lightweight, plug‑and‑play framework that uses a 4 M‑parameter 3D‑Gaussian canonicalizer to adapt frozen Vision‑Language‑Action (VLA) policies to viewpoint shifts without retraining the policy. By treating viewpoint changes as a localized novel‑view synthesis problem under a locality assumption, the method normalizes observations through a scene‑ and policy‑independent disocclusion task. Experiments on the LIBERO benchmark demonstrate that GS‑VLA recovers a large portion of performance lost due to camera displacement, improving results across different policy architectures, unseen task suites, and perturbation scales.
whyItMatters":"The approach offers a computationally efficient alternative to costly fine‑tuning or generative augmentation, enabling robust VLA deployment in real‑world settings where camera configurations may vary."
By Yechan Park, HyunJin Kim
IVSGround introduces a lightweight view selector that learns to choose the most informative camera views for vision‑language model (VLM) based 3D visual grounding, replacing heuristic view selection. The selector is trained via a two‑stage rejection sampling process that uses feedback from a reasoning VLM to generate supervision signals. Experiments on ScanRefer and NR3D demonstrate that IVSGround consistently improves grounding accuracy over existing zero‑shot pipelines, underscoring the importance of selecting where to look for effective 3D visual grounding.
By Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu, Chun-Wei Huang, Quan Kong, Chun-Yi Lee
The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.
By Adam D. Hines, Michael Milford, Tobias Fischer
arXiv:2603.26788v3 Announce Type: replace-cross
Abstract: Zero-shot object navigation requires agents to locate unseen targets in unfamiliar environments without prior maps or task-specific training....
By Feng Wu, Wei Zuo, Wenliang Yang, Jun Xiao, Yang Liu, Xinhua Zeng
arXiv:2606. 04111v1 Announce Type: cross Abstract: Indoor UAV navigation requires efficient exploration, scene understanding, and reliable trajectory execution under limited field-of-view observations.
By Faryal Batool, Muhammad Ahsan Mustafa, Fawad Mehboob, Valerii Serpiva, Dzmitry Tsetserukou
The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.
By Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan
UniGeo is a multimodal large language model designed for text-guided drone geo‑localization, enabling the identification of target regions in large image galleries from natural‑language descriptions. It integrates geo‑semantic understanding, cross‑view semantic generation, and candidate‑level verification within a shared vision‑language framework, establishing stable correspondences among local scene elements, spatial relations, and language. A multi‑stage training strategy progressively refines geo‑semantic learning, cross‑view mapping, and fine‑grained verification, yielding significant performance gains on GeoText‑1652, with R@10 and mAP improvements of 13.59 and 2.83 percentage points respectively.
By Jiahao Wen, Hang Yu, Zhedong Zheng
arXiv:2505.07622v2 Announce Type: replace
Abstract: Cross-view geo-localization is a promising solution for large-scale localization problems, requiring the sequential execution of retrieval and metr...
By Zhuo Song, Ye Zhang, Kunhong Li, Longguang Wang, Yulan Guo