arXiv Computer Vision

Towards Active Cross-View Object Geo-Localization

The paper introduces Active Cross-View Object Geo-Localization (ActiveGeo), enabling mobile agents to actively select new viewpoints and decide when to stop to improve localization with fewer observations. It proposes the ActiveMoPT framework, which uses a three-stage training process: Multi-View Prompt-Preserving Adaptation, Trajectory-Guided Policy Initialization, and Cost-Aware Policy Refinement with GRPO. The authors also create a zero-shot test set, ActiveGeo-858, and demonstrate that ActiveMoPT outperforms prior methods on MoP-UAV and ActiveGeo-858.

arXiv AI
Aug 20

GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting

GS‑VLA introduces a lightweight, plug‑and‑play framework that uses a 4 M‑parameter 3D‑Gaussian canonicalizer to adapt frozen Vision‑Language‑Action (VLA) policies to viewpoint shifts without retraining the policy. By treating viewpoint changes as a localized novel‑view synthesis problem under a locality assumption, the method normalizes observations through a scene‑ and policy‑independent disocclusion task. Experiments on the LIBERO benchmark demonstrate that GS‑VLA recovers a large portion of performance lost due to camera displacement, improving results across different policy architectures, unseen task suites, and perturbation scales. whyItMatters":"The approach offers a computationally efficient alternative to costly fine‑tuning or generative augmentation, enabling robust VLA deployment in real‑world settings where camera configurations may vary."

By Yechan Park, HyunJin Kim
arXiv Computer Vision
Sep 7

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

IVSGround introduces a lightweight view selector that learns to choose the most informative camera views for vision‑language model (VLM) based 3D visual grounding, replacing heuristic view selection. The selector is trained via a two‑stage rejection sampling process that uses feedback from a reasoning VLM to generate supervision signals. Experiments on ScanRefer and NR3D demonstrate that IVSGround consistently improves grounding accuracy over existing zero‑shot pipelines, underscoring the importance of selecting where to look for effective 3D visual grounding.

By Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu, Chun-Wei Huang, Quan Kong, Chun-Yi Lee
arXiv Computer Vision
Sep 21

Multi-viewpoint Geo-localization with Event Cameras

The paper presents MegaEvent, an event‑based visual place recognition system that remains robust to viewpoint changes. By converting five large‑scale geo‑tagged datasets into synthetic event streams and fine‑tuning a vision transformer with a multi‑loss function, MegaEvent achieves an average Recall@1 of 82% on three event‑based localization datasets, outperforming existing methods by 20 recall points. The authors also introduce the Springfield‑Event‑VPR dataset, a 3.7 km walking route recorded in three camera orientations, where MegaEvent surpasses the strongest baseline by 9 recall points.

By Adam D. Hines, Michael Milford, Tobias Fischer
arXiv Computer Vision
3d ago

Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding

The paper introduces CamVLM, a framework that equips large vision‑language models with the ability to actively control camera viewpoints for improved surveillance video understanding. It presents two new datasets: CCTV‑Anomaly, a large‑scale surveillance video collection with detailed captions and event annotations, and CamTrack‑53K, an object‑centric viewpoint trajectory dataset for learning camera actions. Using reinforcement learning, CamVLM learns long‑horizon observation strategies, achieving state‑of‑the‑art performance in both passive and dynamic viewpoint settings.

By Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu, Chen Qian, Shichao Kan
arXiv Computer Vision
Aug 28

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

UniGeo is a multimodal large language model designed for text-guided drone geo‑localization, enabling the identification of target regions in large image galleries from natural‑language descriptions. It integrates geo‑semantic understanding, cross‑view semantic generation, and candidate‑level verification within a shared vision‑language framework, establishing stable correspondences among local scene elements, spatial relations, and language. A multi‑stage training strategy progressively refines geo‑semantic learning, cross‑view mapping, and fine‑grained verification, yielding significant performance gains on GeoText‑1652, with R@10 and mAP improvements of 13.59 and 2.83 percentage points respectively.

By Jiahao Wen, Hang Yu, Zhedong Zheng